From Sights to Insights: Towards Summarization of Multimodal Clinical Documents
Akash Ghosh, Mohit Tomar, Abhisek Tiwari, Sriparna Saha, Jatin Salve, Setu Sinha
摘要
The advancement of Artificial Intelligence is pivotal in reshaping healthcare, enhancing diagnostic precision, and facilitating personalized treatment strategies. One major challenge for healthcare professionals is quickly navigating through long clinical documents to provide timely and effective solutions. Doctors often struggle to draw quick conclusions from these extensive documents. To address this issue and save time for healthcare professionals, an effective summarization model is essential. Most current models assume the data is only textbased. However, patients often include images of their medical conditions in clinical documents. To effectively summarize these multimodal documents, we introduce EDI-Summ, an innovative Image-Guided Encoder-Decoder Model. This model uses modality-aware contextual attention on the encoder and an image cross-attention mechanism on the decoder, enhancing the BART base model to create detailed visual-guided summaries. We have tested our model extensively on three multimodal clinical benchmarks involving multimodal question and dialogue summarization tasks. Our analysis demonstrates that EDI-Summ outperforms state-of-the-art large language and vision-aware models in these summarization tasks. Disclaimer: The work includes vivid medical illustrations, depicting the essential aspects of the subject matter.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- CLINIC : Evaluating Multilingual Trustworthiness in Language Models for HealthcareAkash Ghosh, Srivarshinee Sridhar, Raghav Kaushik Ravi, Muhsin Muhsin 等ICML 2026 · 被引用 6 次
- DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian CultureArijit Maji, Raghvendra Kumar, Akash Ghosh, Anushka 等EMNLP 2025 · 被引用 1 次
- Infogen: Generating Complex Statistical Infographics from DocumentsAkash Ghosh, Aparna Garimella, Pritika Ramu, Sambaran Bandyopadhyay 等ACL 2025
- When Background Matters: Breaking Medical Vision Language Models by Transferable AttackAkash Ghosh, Subhadip Baidya, Sriparna Saha, Xiuying ChenACL 2026
它引用的顶会 Paper7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- A Multitask Framework for Sentiment, Emotion and Sarcasm aware Cyberbullying Detection from Multi-modal Code-Mixed MemesKrishanu Maity, Prince Jha, Sriparna Saha, Pushpak BhattacharyyaSIGIR 2022 · 被引用 80 次
- When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party DialoguesShivani Kumar, Atharva Kulkarni, Md. Shad Akhtar, Tanmoy ChakrabortyACL 2022 · 被引用 54 次
相关 Paper
- UniMS: A Unified Framework for Multimodal Summarization with Knowledge DistillationZhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang 等AAAI 2022 · 被引用 61 次
- Adapting Generative Pretrained Language Model for Open-domain Multimodal Sentence SummarizationDengtian Lin, Liqiang Jing, Xuemeng Song, Meng Liu 等SIGIR 2023 · 被引用 15 次
- mRedditSum: A Multimodal Abstractive Summarization Dataset of Reddit Threads with ImagesKeighley Overbay, Jaewoo Ahn, Fatemeh Pesaran Zadeh, Joonsuk Park 等EMNLP 2023 · 被引用 5 次
- FaMeSumm: Investigating and Improving Faithfulness of Medical SummarizationNan Zhang, Yusen Zhang, Wu Guo, Prasenjit Mitra 等EMNLP 2023 · 被引用 8 次
- Fine-tuning with Multi-modal Entity Prompts for News Image CaptioningJingjing Zhang, Shancheng Fang, Zhendong Mao, Zhiwei Zhang 等ACM MM 2022 · 被引用 16 次
