Cross-Counter-Repeat Attention for Enhanced Understanding of Visual Semantics in Radiology Report Generation
Xiaolei Bo, Feiyang Yang, Feilong Xu, Xiaoli Zhang
摘要
Radiology report generation (RRG), intended to automatically generate a coherent free-text report describing the clinical observations of a radiograph, has been attracting increasing attention from researchers. In recent years, the Transformer-based encoder-decoder architecture has been adopted by most existing methods. However, they neglect the structural rationality issue when applying this single-modal architecture to the multi-modal RRG task, where information can only flow from visual features to textual features, but not in the opposite direction. This information asymmetry results in visual features having no knowledge of the textual features, sending out all visual information, including a large amount of heterogeneous noise. Consequently, this introduces significant resistance to the downstream decoder, which substantially limits or even harms the generation process. To tackle this problem, we present a method where a cross-counter-repeat attention is developed to integrate useful information from two separate modalities, and a memory-driven visual semantics enhancing module is designed to reinforce the visual features with strong time-ordered semantic information. Experimental results on the widely-used IU-Xray dataset show that our approach achieves the state-of-the-art performance, with a remarkable 6.9% improvement in BLEU-4 score. Further analyses also demonstrate that our method can generate sufficiently comprehensive reports to assist radiologists in their clinical decision-making.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- KiUT: Knowledge-injected U-Transformer for Radiology Report GenerationZhongzhen Huang, Xiaofan Zhang, Shaoting ZhangCVPR 2023
- Automatic Radiology Reports Generation via Memory Alignment NetworkHongyu Shen, Mingtao Pei, Juncai Liu, Zhaoxing TianAAAI 2024 · 被引用 40 次
- Unify, Align and Refine: Multi-Level Semantic Alignment for Radiology Report GenerationYaowei Li, Bang Yang, Xuxin Cheng, Zhihong Zhu 等ICCV 2023 · 被引用 47 次
- Visual-Textual Attentive Semantic Consistency for Medical Report GenerationYi Zhou, Lei Huang, Tao Zhou, Huazhu Fu 等ICCV 2021 · 被引用 27 次
- Cross-modal Memory Networks for Radiology Report GenerationZhihong Chen, Yaling Shen, Yan Song, Xiang WanACL 2021
