Mitigating Entity Hallucinations in 3D Radiology Report Generation via Dual-Stream Alignment
Lingyu Zhou, Yue Yu, Zhang Yi, Xiuyuan Xu
摘要
Entity hallucination poses a major challenge in radiology report generation (RRG), particularly for 3D CT scans where complex spatial contexts amplify factual errors. To address this, medical entity phrases serve as key carriers for multi-modal prompting, integrating expert knowledge into the vision-language model. Current methods use unified cross-attention for volume-phrase alignment, failing to account for anatomical specificity during the alignment process. In this work, we introduce the Dual-stream Entity Alignment Reporting network (DEAR) that separately models organ and lesion entities to resolve anatomical bias. Specifically, the dual-stream entity aligner is designed to partition medical entity phrases into organ and lesion streams, feeding them into separate cross-attention blocks in parallel to achieve fine-grained volume–phrase alignment. For structurally regular and spatially stable organ entities, an organ-guided cross-attention (OGCA) block is proposed to enforce structural consistency by retrieving the top-k voxel tokens via volume–phrase similarity and preserving spatial connectivity through morphological dilation. Meanwhile, a lesion-guided cross-attention (LGCA) block is introduced for structurally irregular and spatially variable lesion entities, enhancing anomaly sensitivity through phrase-weighted attention and refining discriminative boundaries via 3D residual Laplacian filtering. Experiments demonstrate that DEAR significantly reduces entity hallucinations and improves clinical factuality in 3D RRG benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Generating Radiology Reports via Memory-driven TransformerZhihong Chen, Yan Song, Tsung-Hui Chang, Xiang WanEMNLP 2020 · 被引用 552 次
- Bootstrapping Large Language Models for Radiology Report GenerationChang Liu, Yuanhe Tian, Weidong Chen, Yan Song 等AAAI 2024 · 被引用 84 次
- CAT: Coordinating Anatomical-Textual Prompts for Multi-Organ and Tumor SegmentationZhongzhen Huang, Yankai Jiang, Rongzhao Zhang, Shaoting Zhang 等NeurIPS 2024 · 被引用 24 次
- Radiology Report Generation via Multi-objective Preference OptimizationTing Xiao, Lei Shi, Peng Liu, Zhe Wang 等AAAI 2025 · 被引用 21 次
相关 Paper
- MEPNet: Medical Entity-Balanced Prompting Network for Brain CT Report GenerationXiaodan Zhang, Yanzhao Shi, Junzhong Ji, Chengxin Zheng 等AAAI 2025 · 被引用 5 次
- DiA-gnostic VLVAE: Disentangled Alignment-Constrained Vision Language Variational AutoEncoder for Robust Radiology Reporting with Missing ModalitiesNagur Shareef Shaik, Teja Krishna Cherukuri, Adnan Masood, Dong Hye YeAAAI 2026
- S2D-Align: Shallow-to-Deep Auxiliary Learning for Anatomically-Grounded Radiology Report GenerationJiechao Gao, Chang Liu, Yuangang LiAAAI 2026
- Rethinking Radiology Report Generation: From Narrative Flow to Topic-Guided FindingsSheng Cheng, Devika SubramanianICLR 2026
- A Disease-Aware Dual-Stage Framework for Chest X-ray Report GenerationPuzhen Wu, Hexin Dong, Yi Lin, Yihao Ding 等AAAI 2026 · 被引用 3 次
