Unify, Align and Refine: Multi-Level Semantic Alignment for Radiology Report Generation
Yaowei Li, Bang Yang, Xuxin Cheng, Zhihong Zhu, Hongxiang Li, Yuexian Zou
摘要
Automatic radiology report generation has attracted enormous research interest due to its practical value in reducing the workload of radiologists. However, simultaneously establishing global correspondences between the image (e.g., Chest X-ray) and its related report and local alignments between image patches and keywords remains challenging. To this end, we propose an Unify, Align and then Refine (UAR) approach to learn multi-level cross-modal alignments and introduce three novel modules: Latent Space Unifier (LSU), Cross-modal Representation Aligner (CRA) and Text-to-Image Refiner (TIR). Specifically, LSU unifies multimodal data into discrete tokens, making it flexible to learn common knowledge among modalities with a shared network. The modality-agnostic CRA learns discriminative features via a set of orthonormal basis and a dual-gate mechanism first and then globally aligns visual and textual representations under a triplet contrastive loss. TIR boosts token-level local alignment via calibrating text-to-image attention with a learnable mask. Additionally, we design a two-stage training procedure to make UAR gradually grasp cross-modal alignments at different levels, which imitates radiologists’ workflow: writing sentence by sentence first and then checking word by word. Extensive experiments and analyses on IU-Xray and MIMIC-CXR benchmark datasets demonstrate the superiority of our UAR against varied state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Implicit Neural Representation for Cooperative Low-light Image EnhancementShuzhou Yang, Moxuan Ding, Yanmin Wu, Zihan Li 等ICCV 2023 · 被引用 224 次
- Grounding 3D Object Affordance from 2D Interactions in ImagesYuhang Yang, Wei Zhai, Hongchen Luo, Yang Cao 等ICCV 2023 · 被引用 69 次
- G2L: Semantically Aligned and Uniform Video Grounding via Geodesic and Game TheoryHongxiang Li, Meng Cao, Xuxin Cheng, Yaowei Li 等ICCV 2023 · 被引用 32 次
- Image-aware Evaluation of Generated Medical ReportsGefen Dawidowicz, Elad Hirsch, Ayellet TalNeurIPS 2024 · 被引用 3 次
- DAMPER: A Dual-Stage Medical Report Generation Framework with Coarse-Grained MeSH Alignment and Fine-Grained Hypergraph MatchingXiaofei Huang, Wenting Chen, Jie Liu, Qisheng Lu 等AAAI 2025 · 被引用 2 次
它引用的顶会 Paper14
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu 等AAAI 2020 · 被引用 1,047 次
- Generating Radiology Reports via Memory-driven TransformerZhihong Chen, Yan Song, Tsung-Hui Chang, Xiang WanEMNLP 2020 · 被引用 552 次
- When Radiology Report Generation Meets Knowledge GraphYixiao Zhang, Xiaosong Wang, Ziyue Xu, Qihang Yu 等AAAI 2020 · 被引用 391 次
相关 Paper
- Automatic Radiology Reports Generation via Memory Alignment NetworkHongyu Shen, Mingtao Pei, Juncai Liu, Zhaoxing TianAAAI 2024 · 被引用 40 次
- Bootstrapping Large Language Models for Radiology Report GenerationChang Liu, Yuanhe Tian, Weidong Chen, Yan Song 等AAAI 2024 · 被引用 84 次
- Visual-Textual Attentive Semantic Consistency for Medical Report GenerationYi Zhou, Lei Huang, Tao Zhou, Huazhu Fu 等ICCV 2021 · 被引用 27 次
- KiUT: Knowledge-injected U-Transformer for Radiology Report GenerationZhongzhen Huang, Xiaofan Zhang, Shaoting ZhangCVPR 2023
- A Disease-Aware Dual-Stage Framework for Chest X-ray Report GenerationPuzhen Wu, Hexin Dong, Yi Lin, Yihao Ding 等AAAI 2026 · 被引用 3 次
