Auto-Encoding Knowledge Graph for Unsupervised Medical Report Generation
Fenglin Liu, Chenyu You, Xian Wu, Shen Ge, Sheng Wang, Xu Sun
Abstract
Medical report generation, which aims to automatically generate a long and coherent report of a given medical image, has been receiving growing research interests. Existing approaches mainly adopt a supervised manner and heavily rely on coupled image-report pairs. However, in the medical domain, building a large-scale image-report paired dataset is both time-consuming and expensive. To relax the dependency on paired data, we propose an unsupervised model Knowledge Graph Auto-Encoder (KGAE) which accepts independent sets of images and reports in training. KGAE consists of a pre-constructed knowledge graph, a knowledge-driven encoder and a knowledge-driven decoder. The knowledge graph works as the shared latent space to bridge the visual and textual domains; The knowledge-driven encoder projects medical images and reports to the corresponding coordinates in this latent space and the knowledge-driven decoder generates a medical report given a coordinate in this space. Since the knowledge-driven encoder and decoder can be trained with independent sets of images and reports, KGAE is unsupervised. The experiments show that the unsupervised KGAE generates desirable medical reports without using any image-report training pairs. Moreover, KGAE can also work in both semi-supervised and supervised settings, and accept paired images and reports in training. By further fine-tuning with image-report pairs, KGAE consistently outperforms the current state-of-the-art models on two datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac0bc088-798f-4cc6-b8d3-62a57c4a30b4Cited by top-tier papers12
- Prophet Attention: Predicting Attention with Future AttentionFenglin Liu, Xuancheng Ren, Xian Wu, Shen Ge et al.NeurIPS 2020 · 52 citations
- EMVLight: A Decentralized Reinforcement Learning Framework for Efficient Passage of Emergency VehiclesHaoran Su, Yaofeng Desmond Zhong, Biswadip Dey, Amit ChakrabortyAAAI 2022 · 29 citations
- DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only TrainingWei Li, Linchao Zhu, Longyin Wen, Yi YangICLR 2023 · 24 citations
- Self-supervised Spatial Reasoning on Multi-View Line DrawingsSiyuan Xiang, Anbang Yang, Yanfei Xue, Yaoqing Yang et al.CVPR 2022 · 4 citations
- Image-aware Evaluation of Generated Medical ReportsGefen Dawidowicz, Elad Hirsch, Ayellet TalNeurIPS 2024 · 3 citations
Builds on3
- Generating Radiology Reports via Memory-driven TransformerZhihong Chen, Yan Song, Tsung-Hui Chang, Xiang WanEMNLP 2020 · 552 citations
- When Radiology Report Generation Meets Knowledge GraphYixiao Zhang, Xiaosong Wang, Ziyue Xu, Qihang Yu et al.AAAI 2020 · 391 citations
- Meshed-Memory Transformer for Image CaptioningMarcella Cornia, Matteo Stefanini, Lorenzo Baraldi, Rita CucchiaraCVPR 2020
Related papers
- DiA-gnostic VLVAE: Disentangled Alignment-Constrained Vision Language Variational AutoEncoder for Robust Radiology Reporting with Missing ModalitiesNagur Shareef Shaik, Teja Krishna Cherukuri, Adnan Masood, Dong Hye YeAAAI 2026
- Divide and Conquer: Isolating Normal-Abnormal Attributes in Knowledge Graph-Enhanced Radiology Report GenerationXiao Liang, Yanlei Zhang, Di Wang, Haodi Zhong et al.ACM MM 2024 · 7 citations
- Unpaired Image Captioning via Scene Graph AlignmentsJiuxiang Gu, Shafiq R. Joty, Jianfei Cai, Handong Zhao et al.ICCV 2019 · 191 citations
- Dynamic Graph Enhanced Contrastive Learning for Chest X-Ray Report GenerationMingjie Li, Bingqian Lin, Zicong Chen, Haokun Lin et al.CVPR 2023
- Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report GenerationLongzhen Yang, Zhangkai Ni, Ying Wen, Yihang Liu et al.ACM MM 2025
