Automatic Radiology Reports Generation via Memory Alignment Network
Hongyu Shen, Mingtao Pei, Juncai Liu, Zhaoxing Tian
Abstract
The automatic generation of radiology reports is of great significance, which can reduce the workload of doctors and improve the accuracy and reliability of medical diagnosis and treatment, and has attracted wide attention in recent years. Cross-modal mapping between images and text, a key component of generating high-quality reports, is challenging due to the lack of corresponding annotations. Despite its importance, previous studies have often overlooked it or lacked adequate designs for this crucial component. In this paper, we propose a method with memory alignment embedding to assist the model in aligning visual and textual features to generate a coherent and informative report. Specifically, we first get the memory alignment embedding by querying the memory matrix, where the query is derived from a combination of the visual features and their corresponding positional embeddings. Then the alignment between the visual and textual features can be guided by the memory alignment embedding during the generation process. The comparison experiments with other alignment methods show that the proposed alignment method is less costly and more effective. The proposed approach achieves better performance than state-of-the-art approaches on two public datasets IU X-Ray and MIMIC-CXR, which further demonstrates the effectiveness of the proposed alignment method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e39b3da7-f8b3-4401-957a-a2b2cc872445Cited by top-tier papers10
- Radiology Report Generation via Multi-objective Preference OptimizationTing Xiao, Lei Shi, Peng Liu, Zhe Wang et al.AAAI 2025 · 21 citations
- HC-LLM: Historical-Constrained Large Language Models for Radiology Report GenerationTengfei Liu, Jiapu Wang, Yongli Hu, Mingjie Li et al.AAAI 2025 · 6 citations
- PriorRG: Prior-Guided Contrastive Pre-training and Coarse-to-Fine Decoding for Chest X-ray Report GenerationKang Liu, Zhuoqi Ma, Zikang Fang, Yunan Li et al.AAAI 2026 · 5 citations
- OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report GenerationZhuoxiao Chen, Hongyang Yu, Ying Xu, Yadan Luo et al.CVPR 2026 · 3 citations
- Online Iterative Self-Alignment for Radiology Report GenerationTing Xiao, Lei Shi, Yang Zhang, HaoFeng Yang et al.ACL 2025 · 2 citations
Builds on6
- Generating Radiology Reports via Memory-driven TransformerZhihong Chen, Yan Song, Tsung-Hui Chang, Xiang WanEMNLP 2020 · 552 citations
- Clinical-BERT: Vision-Language Pre-training for Radiograph Diagnosis and Reports GenerationBin Yan, Mingtao PeiAAAI 2022 · 138 citations
- Learning Distinct and Representative Modes for Image CaptioningQi Chen, Chaorui Deng, Qi WuNeurIPS 2022 · 27 citations
- Transform and Tell: Entity-Aware News Image CaptioningAlasdair Tran, Alexander Patrick Mathews, Lexing XieCVPR 2020
- Cross-modal Memory Networks for Radiology Report GenerationZhihong Chen, Yaling Shen, Yan Song, Xiang WanACL 2021
Related papers
- Unify, Align and Refine: Multi-Level Semantic Alignment for Radiology Report GenerationYaowei Li, Bang Yang, Xuxin Cheng, Zhihong Zhu et al.ICCV 2023 · 47 citations
- Visual-Textual Attentive Semantic Consistency for Medical Report GenerationYi Zhou, Lei Huang, Tao Zhou, Huazhu Fu et al.ICCV 2021 · 27 citations
- S2D-Align: Shallow-to-Deep Auxiliary Learning for Anatomically-Grounded Radiology Report GenerationJiechao Gao, Chang Liu, Yuangang LiAAAI 2026
- Cross-Counter-Repeat Attention for Enhanced Understanding of Visual Semantics in Radiology Report GenerationXiaolei Bo, Feiyang Yang, Feilong Xu, Xiaoli ZhangACM MM 2025
- Learnable Retrieval Enhanced Visual-Text Alignment and Fusion for Radiology Report GenerationQin Zhou, Guoyan Liang, Xindi Li, Jingyuan Chen et al.ICCV 2025 · 2 citations
