Medical Report Generation via Multimodal Spatio-Temporal Fusion
Xin Mei, Rui Mao, Xiaoyan Cai, Libin Yang, Erik Cambria
摘要
Medical report generation aims at automating the synthesis of accurate and comprehensive diagnostic reports from radiological images. The task can significantly enhance clinical decision-making and alleviate the workload on radiologists. Existing works normally generate reports from single chest radiographs, although historical examination data also serve as crucial references for radiologists in real-world clinical settings. To address this constraint, we introduce a novel framework that mimics the workflow of radiologists. This framework compares past and present patient images to monitor disease progression and incorporates prior diagnostic reports as references for generating current personalized reports. We tackle the textual diversity challenge in cross-modal tasks by promoting style-agnostic discrete report representation learning and token generation. Furthermore, we propose a novel spatio-temporal fusion method with multi-granularities to fuse textual and visual features by disentangling the differences between current and historical data. We also tackle token generation biases, which arise from long-tail frequency distributions, proposing a novel feature normalization technique. This technique ensures unbiased generation for tokens, whether they are frequent or infrequent, enabling the robustness of report generation for rare diseases. Experimental results on the two public datasets demonstrate that our proposed model outperforms state-of-the-art baselines.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- PriorRG: Prior-Guided Contrastive Pre-training and Coarse-to-Fine Decoding for Chest X-ray Report GenerationKang Liu, Zhuoqi Ma, Zikang Fang, Yunan Li 等AAAI 2026 · 被引用 5 次
- Personalized Longitudinal Medical Report Generation via Temporally-Aware Federated AdaptationHe Zhu, Ren Togo, Takahiro Ogawa, Kenji Hirata 等CVPR 2026 · 被引用 1 次
- RefleXNet: Targeted Self-Reflection for Accurate Chest X-ray ReportingXin Mei, Rui Mao, Xiaoyan Cai, Libin Yang 等AAAI 2026
- FAMDR: Feature-Aligned Multimodal Denoising for Reliable Diagnostic Reconciliation in Medical ImagingXun Liang, Zhiying Li, Hongxun JiangAAAI 2026
相关 Paper
- Enhanced Contrastive Learning with Multi-view Longitudinal Data for Chest X-ray Report GenerationKang Liu, Zhuoqi Ma, Xiaolu Kang, Yunan Li 等CVPR 2025
- A Disease-Aware Dual-Stage Framework for Chest X-ray Report GenerationPuzhen Wu, Hexin Dong, Yi Lin, Yihao Ding 等AAAI 2026 · 被引用 3 次
- Unify, Align and Refine: Multi-Level Semantic Alignment for Radiology Report GenerationYaowei Li, Bang Yang, Xuxin Cheng, Zhihong Zhu 等ICCV 2023 · 被引用 47 次
- HC-LLM: Historical-Constrained Large Language Models for Radiology Report GenerationTengfei Liu, Jiapu Wang, Yongli Hu, Mingjie Li 等AAAI 2025 · 被引用 6 次
- Automated Generation of Accurate & Fluent Medical X-ray ReportsHoang T. N. Nguyen, Dong Nie, Taivanbat Badamdorj, Yujie Liu 等EMNLP 2021 · 被引用 36 次
