Notes — Multi-modal LLMs for Medical Imaging
Published:
Key Ideas
- Vision-language alignment needs clinical context supervision.
- Reporting-style training targets may improve factual consistency.
Published:
Key Ideas
Published:
Short description of portfolio item number 1
Short description of portfolio item number 2 
Published:
Published in Journal of Physics: Conference Series, 2023
A comparative analysis of LXMERT and UNITER models for visual question answering tasks.
Recommended citation: Zhicheng He, Yuanzhi Li, Dingming Zhang. (2023). "Transformer-based visual question answering model comparison." Journal of Physics: Conference Series, 2646(1), 012031.
Download Paper
Published in Biomedical Engineering and Computational Biology, 2024
Fine-tuning MedSAM for impacted tooth segmentation in X-ray images to aid dental diagnoses.
Recommended citation: He Zhicheng, Wang Yipeng, Li Xiao. (2024). "Deep Learning-Based Detection of Impacted Teeth on Panoramic Radiographs." Biomedical Engineering and Computational Biology, 15, 11795972241288319.
Download Paper
Published in International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2024
A WCE unified illumination correction solution using an end-to-end promptable diffusion transformer (DiT) model.
Recommended citation: Long Bai, Tong Chen, Qiaozhi Tan, Wan Jun Nah, Yanheng Li, Zhicheng He, Sishen Yuan, Zhen Chen, Jinlin Wu, Mobarakol Islam, Zhen Li, Hongbin Liu, Hongliang Ren. (2024). "Endouic: Promptable diffusion transformer for unified illumination correction in capsule endoscopy." International Conference on Medical Image Computing and Computer-Assisted Intervention, Pages 296-306.
Download Paper
Published in 2024 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), 2024
A graph self-supervised learning framework utilizing graph matching for molecular property prediction.
Recommended citation: Hongxiang Lin, Yixiao Zhou, Huiying Hu, Zhicheng He, Runzhi Wu, Xiaoqing Lyu. (2024). "Graph Matching Based Graph Self-Supervised Learning for Molecular Property Prediction." 2024 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), Pages 7092-7094.
Download Paper
Published in International Conference on Document Analysis and Recognition (ICDAR), 2025
Paper2PPT framework for generating visual presentations of scientific papers with cross-modal alignment.
Recommended citation: Huiying Hu, Zhicheng He, Yixiao Zhou, Tongwei Zhang, Xiaoqing Lyu. (2025). "Multimodal Content Alignment with LLM for Visual Presentation of Papers." International Conference on Document Analysis and Recognition, Pages 238-256.
Download Paper
Published in International Conference on Document Analysis and Recognition (ICDAR), 2025
A self-prompted segmentation framework for scientific illustrations using SAM-based methods.
Recommended citation: Tuo Wang, Yixiao Zhou, Tongwei Zhang, Zhicheng He, Yumeng Zhao, Xiaoqing Lyu. (2025). "SSSI: Self-prompted Segmentation of Scientific Illustrations." International Conference on Document Analysis and Recognition, Pages 347-361.
Download Paper
Published in Medical Imaging with Deep Learning (MIDL), 2026
A task-oriented feature disentanglement framework (DINOv3-FD) for parameter-efficient adaptation of DINOv3 to medical vision tasks.
Recommended citation: Zhicheng He, Yibing Fu, Yueming Jin. (2026). "Incentivizing DINOv3 Adaptation for Medical Vision Tasks via Feature Disentanglement." Medical Imaging with Deep Learning.
Download Paper
Published in arXiv preprint arXiv:2602.14512 (v3), 2026
A next-scale autoregressive foundation model for CT and MRI generation across six anatomical regions.
Recommended citation: Zhicheng He, Yunpeng Zhao, Junde Wu, Ziwei Niu, Ziyue Wang, Bohan Li, Zijun Li, Lanfen Lin, Nan Liu, Yueming Jin. (2026). "Scalable next-scale autoregression for medical image generation across anatomical regions." arXiv preprint arXiv:2602.14512 (v3).
Download Paper
Published in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
A multimodal model that adapts visual feature granularity to the text context.
Recommended citation: Junyuan Mao, Qiankun Li, Linghao Meng, Zhicheng He, Xinliang Zhou, Kun Wang, Yang Liu, Yueming Jin. (2026). "Granulon: Awakening Pixel-Level Visual Encoders with Adaptive Multi-Granularity Semantics for MLLM." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Pages 26317-26327.
Download Paper
Published in International Workshop on Document Analysis Systems (DAS), 2026
Structure-aware representation learning for understanding scientific method diagrams.
Recommended citation: Zhicheng He, Zheyi Deng, Junhao Ren, Xiaoqing Lyu. (2026). "SciXplain: Structure-Aware Representation Learning for Scientific Method Diagrams." Document Analysis Systems (DAS 2026), Pages 474-490.
Download Paper
Published in Biocybernetics and Biomedical Engineering, 2026
A CBCT-based graph framework for skeletal classification of malocclusion.
Recommended citation: Zhichun Jin*, Zhicheng He*, Hao Xu, Dongyang Li, Lin Wang, Hongliang Ren, Long Bai. (2026). "Morphological Decoupling-Based Skeletal Classification for Clinical Assessment of Malocclusion." Biocybernetics and Biomedical Engineering, 46(4), Pages 650-660.
Download Paper
Published in European Conference on Computer Vision (ECCV), 2026
A memory-guided keyframe selection framework for long-video question answering.
Recommended citation: Linghao Meng, Qiankun Li, Junyuan Mao, Pujin Liao, Zhicheng He, Enbo Zhang, Kun Wang, Yang Liu, Huazhu Fu, Yueming Jin. (2026). "Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding." European Conference on Computer Vision, Pages 588-606.
Download Paper
Published:
This is a description of your talk, which is a markdown files that can be all markdown-ified like any other post. Yay markdown!
Published:
This is a description of your conference proceedings talk, note the different field in type. You can put anything in this field.
Undergraduate course, University 1, Department, 2014
This is a description of a teaching experience. You can use markdown like any other post.
Workshop, University 1, Department, 2015
This is a description of a teaching experience. You can use markdown like any other post.