Visual-Textual Attentive Semantic Consistency for Medical Report Generation
Yi Zhou, Lei Huang, Tao Zhou, Huazhu Fu, Ling Shao
摘要
Automatic report generation on medical radiographs have recently gained interest. However, identifying diseases as well as correctly predicting their corresponding sizes, locations and other medical description patterns, which is essential for generating high-quality reports, is challenging. Although previous methods focused on producing readable reports, how to accurately detect and describe findings that match with the query X-Ray has not been successfully addressed. In this paper, we propose a multi-modality semantic attention model to integrate visual features, predicted key finding embeddings, as well as clinical features, and progressively decode reports with visual-textual semantic consistency. First, multi-modality features are extracted and attended with the hidden states from the sentence de-coder, to encode enriched context vectors for better decoding a report. These modalities include regional visual features of scans, semantic word embeddings of the top-K findings predicted with high probabilities, and clinical features of indications. Second, the progressive report decoder consists of a sentence decoder and a word decoder, where we propose image-sentence matching and description accuracy losses to constrain the visual-textual semantic consistency. Extensive experiments on the public MIMIC-CXR and IU X-Ray datasets show that our model achieves consistent improvements over the state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- MMTN: Multi-Modal Memory Transformer Network for Image-Report Consistent Medical Report GenerationYiming Cao, Lizhen Cui, Lei Zhang, Fuqiang Yu 等AAAI 2023 · 被引用 56 次
- Beyond N-grams: A Hierarchical Reward Learning Framework for Clinically-Aware Medical Report GenerationYuan Wang, Shujian Gao, Jiaxiang Liu, Songtao Jiang 等AAAI 2026 · 被引用 2 次
- F-Assist: Multi-Phase Fetal Growth Forecast and Report Generation from Ultrasound ExaminationBin Pu, Xusheng Liang, Xinpeng Ding, Jinlin Wu 等CVPR 2026
它引用的顶会 Paper7
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 被引用 992 次
- When Radiology Report Generation Meets Knowledge GraphYixiao Zhang, Xiaosong Wang, Ziyue Xu, Qihang Yu 等AAAI 2020 · 被引用 391 次
- Entangled Transformer for Image CaptioningGuang Li, Linchao Zhu, Ping Liu, Yi YangICCV 2019 · 被引用 346 次
- Unpaired Image Captioning via Scene Graph AlignmentsJiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao 等ICCV 2019 · 被引用 191 次
- Hierarchy Parsing for Image CaptioningTing Yao, Yingwei Pan, Yehao Li, Tao MeiICCV 2019 · 被引用 183 次
相关 Paper
- Automatic Radiology Reports Generation via Memory Alignment NetworkHongyu Shen, Mingtao Pei, Juncai Liu, Zhaoxing TianAAAI 2024 · 被引用 40 次
- Unify, Align and Refine: Multi-Level Semantic Alignment for Radiology Report GenerationYaowei Li, Bang Yang, Xuxin Cheng, Zhihong Zhu 等ICCV 2023 · 被引用 47 次
- KiUT: Knowledge-injected U-Transformer for Radiology Report GenerationZhongzhen Huang, Xiaofan Zhang, Shaoting ZhangCVPR 2023
- The Impact of Auxiliary Patient Data on Automated Chest X-Ray Report Generation and How to Incorporate ItAaron Nicolson, Shengyao Zhuang, Jason Dowling, Bevan KoopmanACL 2025 · 被引用 6 次
- A Disease-Aware Dual-Stage Framework for Chest X-ray Report GenerationPuzhen Wu, Hexin Dong, Yi Lin, Yihao Ding 等AAAI 2026 · 被引用 3 次
