MMTN: Multi-Modal Memory Transformer Network for Image-Report Consistent Medical Report Generation
Yiming Cao, Lizhen Cui, Lei Zhang, Fuqiang Yu, Zhen Li, Yonghui Xu
Abstract
Automatic medical report generation is an essential task in applying artificial intelligence to the medical domain, which can lighten the workloads of doctors and promote clinical automation. The state-of-the-art approaches employ Transformer-based encoder-decoder architectures to generate reports for medical images. However, they do not fully explore the relationships between multi-modal medical data, and generate inaccurate and inconsistent reports. To address these issues, this paper proposes a Multi-modal Memory Transformer Network (MMTN) to cope with multi-modal medical data for generating image-report consistent medical reports. On the one hand, MMTN reduces the occurrence of image-report inconsistencies by designing a unique encoder to associate and memorize the relationship between medical images and medical terminologies. On the other hand, MMTN utilizes the cross-modal complementarity of the medical vision and language for the word prediction, which further enhances the accuracy of generating medical reports. Extensive experiments on three real datasets show that MMTN achieves significant effectiveness over state-of-the-art approaches on both automatic metrics and human evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f3f471b5-6bdf-4350-b690-0a6aaff30ab3Cited by top-tier papers9
- Radiology Report Generation via Multi-objective Preference OptimizationTing Xiao, Lei Shi, Peng Liu, Zhe Wang et al.AAAI 2025 · 21 citations
- HC-LLM: Historical-Constrained Large Language Models for Radiology Report GenerationTengfei Liu, Jiapu Wang, Yongli Hu, Mingjie Li et al.AAAI 2025 · 6 citations
- Online Iterative Self-Alignment for Radiology Report GenerationTing Xiao, Lei Shi, Yang Zhang, HaoFeng Yang et al.ACL 2025 · 2 citations
- CA-MLIF: Cross-Attention and Multimodal Low-Rank Interaction Fusion Framework for Tumor Prognostic PredictionYajun An, Jiale Chen, Huan Lin, Zhenbing Liu et al.AAAI 2025 · 1 citation
- Cross-Modal Alignment via Variational Copula ModellingFeng Wu, Tsai Hor Chan, Fuying Wang, Guosheng Yin et al.ICML 2025
Builds on7
- Generating Radiology Reports via Memory-driven TransformerZhihong Chen, Yan Song, Tsung-Hui Chang, Xiang WanEMNLP 2020 · 552 citations
- When Radiology Report Generation Meets Knowledge GraphYixiao Zhang, Xiaosong Wang, Ziyue Xu, Qihang Yu et al.AAAI 2020 · 391 citations
- Visual-Textual Attentive Semantic Consistency for Medical Report GenerationYi Zhou, Lei Huang, Tao Zhou, Huazhu Fu et al.ICCV 2021 · 27 citations
- Cross-modal Memory Networks for Radiology Report GenerationZhihong Chen, Yaling Shen, Yan Song, Xiang WanACL 2021
- RSTNet: Captioning With Adaptive Attention on Visual and Non-Visual WordsXuying Zhang, Xiaoshuai Sun, Yunpeng Luo, Jiayi Ji et al.CVPR 2021
Related papers
- KiUT: Knowledge-injected U-Transformer for Radiology Report GenerationZhongzhen Huang, Xiaofan Zhang, Shaoting ZhangCVPR 2023
- Cross-Counter-Repeat Attention for Enhanced Understanding of Visual Semantics in Radiology Report GenerationXiaolei Bo, Feiyang Yang, Feilong Xu, Xiaoli ZhangACM MM 2025
- MMT: Image-guided Story Ending Generation with Multimodal Memory TransformerDizhan Xue, Shengsheng Qian, Quan Fang, Changsheng XuACM MM 2022 · 16 citations
- Automatic Radiology Reports Generation via Memory Alignment NetworkHongyu Shen, Mingtao Pei, Juncai Liu, Zhaoxing TianAAAI 2024 · 40 citations
- MedM2G: Unifying Medical Multi-Modal Generation via Cross-Guided Diffusion with Visual InvariantChenlu Zhan, Yu Lin, Gaoang Wang, Hongwei Wang et al.CVPR 2024 · 20 citations
