Medical Vision-Language Pretraining with LLM-Guided Temporal Supervision
Liang Bai, Zhi Wang, Huimin Yan, Xian Yang
Abstract
Medical vision–language pretraining typically relies on static image–text pairs, overlooking temporal cues vital for understanding clinical progression. This limits model sensitivity to evolving semantics and reduces their effectiveness in real-world clinical reasoning. To address this challenge, we propose TAMM—a temporal alignment framework that leverages weak but semantically rich supervision from large language models (LLMs). Given temporally adjacent clinical reports, LLMs automatically generate (i) coarse-grained trend labels (e.g., improving or worsening), and (ii) fine-grained rationales explaining the supporting clinical evidence. These complementary signals inject temporal semantics without requiring manual annotation, and guide vision–language representation learning to capture trend-sensitive cross-modal alignment and rationale-grounded coherence. Experiments on multiple medical benchmarks demonstrate that TAMM improves retrieval and classification performance while yielding more interpretable, temporally consistent embeddings. Our results highlight the potential of leveraging LLM-derived supervision to equip vision–language models with temporal awareness critical for clinical applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e1ad6735-8405-411f-82d8-8c6eb67949a4Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MedCLIP: Contrastive Learning from Unpaired Medical Images and TextZifeng Wang, Zhenbang Wu, Dinesh Agarwal, Jimeng SunEMNLP 2022 · 907 citations
- GLoRIA: A Multimodal Global-Local Representation Learning Framework for Label-efficient Medical Image RecognitionShih-Cheng Huang, Liyue Shen, Matthew P. Lungren, Serena YeungICCV 2021 · 516 citations
- Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation LearningFuying Wang, Yuyin Zhou, Shujun Wang, Varut Vardhanabhuti et al.NeurIPS 2022 · 302 citations
Related papers
- LLM-Guided Diagnostic Evidence Alignment for Medical Vision–Language Pretraining under Limited PairingHuimin Yan, Liang Bai, Xian Yang, Long ChenICML 2026 · 1 citation
- Knowledge-Augmented Multimodal Clinical Rationale Generation for Disease Diagnosis with Small Language ModelsShuai Niu, Jing Ma, Hongzhan Lin, Liang Bai et al.ACL 2025 · 5 citations
- Temporal Inversion for Learning Interval Change in Chest X-RaysHanbin Ko, Kyeongmin Jeon, Doowoong Choi, Chang Min ParkCVPR 2026 · 3 citations
- Learning to Exploit Temporal Structure for Biomedical Vision-Language ProcessingShruthi Bannur, Stephanie L. Hyland, Qianchu Liu, Fernando Pérez-García et al.CVPR 2023
- PRIOR: Prototype Representation Joint Learning from Medical Images and ReportsPujin Cheng, Li Lin, Junyan Lyu, Yijin Huang et al.ICCV 2023 · 91 citations
