Learnable Retrieval Enhanced Visual-Text Alignment and Fusion for Radiology Report Generation
Qin Zhou, Guoyan Liang, Xindi Li, Jingyuan Chen, Zhe Wang, Chang Yao, Sai Wu
摘要
Automated radiology report generation is essential for improving diagnostic efficiency and reducing the workload of medical professionals. However, existing methods face significant challenges, such as disease class imbalance and insufficient cross-modal fusion. To address these issues, we propose the learnable Retrieval Enhanced Visual-Text Alignment and Fusion (REVTAF) framework, which effectively tackles both class imbalance and visual-text fusion in report generation. REVTAF incorporates two core components: (1) a Learnable Retrieval Enhancer (LRE) that utilizes semantic hierarchies from hyperbolic space and intra-batch context through a ranking-based metric. LRE adaptively retrieves the most relevant reference reports, enhancing image representations, particularly for underrepresented (tail) class inputs; and (2) a fine-grained visual-text alignment and fusion strategy that ensures consistency across multi-source cross-attention maps for precise alignment. This component further employs an optimal transport-based cross-attention mechanism to dynamically integrate task-relevant textual knowledge for improved report generation. By combining adaptive retrieval with multi-source alignment and fusion, REVTAF achieves fine-grained visual-text integration under weak image-report level supervision while effectively mitigating data imbalance issues. The experiments demonstrate that REVTAF outperforms state-of-the-art methods, achieving an average improvement of 7.4% on the MIMIC-CXR dataset and 2.9% on the IU X-Ray dataset. Comparisons with mainstream multimodal LLMs (e.g., GPT-series models), further highlight its superiority in radiology report generation https://github.com/banbooliang/REVTAF-RRG.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Enhancing Reinforcement Learning for Radiology Report Generation with Evidence-aware Rewards and Self-correcting Preference LearningQin Zhou, Guoyan Liang, Qianyi Yang, Jingyuan Chen 等ACL 2026 · 被引用 1 次
- The Double Dilemma in Multi-Task Radiology Report Generation: A Gradient Dynamics Analysis and SolutionErjian Zhang, Yatong Hao, Liejun Wang, Zhiqing GuoICML 2026
- Scalable Medical Multimodal Fusion via Symmetric Consistency ModelingXiaowen Sun, Hui Liu, Gongguan Chen, Ning MaoICML 2026
它引用的顶会 Paper15
- MedCLIP: Contrastive Learning from Unpaired Medical Images and TextZifeng Wang, Zhenbang Wu, Dinesh Agarwal, Jimeng SunEMNLP 2022 · 被引用 907 次
- Generating Radiology Reports via Memory-driven TransformerZhihong Chen, Yan Song, Tsung-Hui Chang, Xiang WanEMNLP 2020 · 被引用 552 次
- Scaling Up Vision-Language Pretraining for Image CaptioningXiaowei Hu, Zhe Gan, Jianfeng Wang, Zhengyuan Yang 等CVPR 2022 · 被引用 203 次
- PromptMRG: Diagnosis-Driven Prompts for Medical Report GenerationHaibo Jin, Haoxuan Che, Yi Lin, Hao ChenAAAI 2024 · 被引用 168 次
- Clinical-BERT: Vision-Language Pre-training for Radiograph Diagnosis and Reports GenerationBin Yan, Mingtao PeiAAAI 2022 · 被引用 138 次
相关 Paper
- Automatic Radiology Reports Generation via Memory Alignment NetworkHongyu Shen, Mingtao Pei, Juncai Liu, Zhaoxing TianAAAI 2024 · 被引用 40 次
- Unify, Align and Refine: Multi-Level Semantic Alignment for Radiology Report GenerationYaowei Li, Bang Yang, Xuxin Cheng, Zhihong Zhu 等ICCV 2023 · 被引用 47 次
- S2D-Align: Shallow-to-Deep Auxiliary Learning for Anatomically-Grounded Radiology Report GenerationJiechao Gao, Chang Liu, Yuangang LiAAAI 2026
- A Disease-Aware Dual-Stage Framework for Chest X-ray Report GenerationPuzhen Wu, Hexin Dong, Yi Lin, Yihao Ding 等AAAI 2026 · 被引用 3 次
- DART: Disease-aware Image-Text Alignment and Self-correcting Re-alignment for Trustworthy Radiology Report GenerationSang-Jun Park, Keun-Soo Heo, Dong-Hee Shin, Young-Han Son 等CVPR 2025
