MEDSAGE: Enhancing Robustness of Medical Dialogue Summarization to ASR Errors with LLM-generated Synthetic Dialogues
Kuluhan Binici, Abhinav Ramesh Kashyap, Viktor Schlegel, Andy T. Liu, Vijay Prakash Dwivedi, Thanh-Tung Nguyen, Xiaoxue Gao, Nancy F. Chen, Stefan Winkler
Abstract
Automatic Speech Recognition (ASR) systems are used to transcribe speech into text, yet the errors they introduce can significantly degrade the performance of downstream tasks like summarization. This issue is particularly pronounced in clinical dialogue summarization, a low-resource domain where supervised data for fine-tuning is scarce, necessitating the use of ASR models as black-box solutions. Employing conventional data augmentation for enhancing the noise robustness of summarization models is not feasible either due to the unavailability of sufficient medical dialogue audio recordings and corresponding ASR transcripts. To address this challenge, we propose MEDSAGE, an approach for generating synthetic samples for data augmentation using Large Language Models (LLMs). Specifically, we leverage the incontext learning capabilities of LLMs and instruct them to generate ASR-like errors based on a few available medical dialogue examples with audio recordings. Experimental results show that LLMs can effectively model ASR noise, and incorporating this noisy data into the training process significantly improves the robustness and accuracy of medical dialogue summarization systems. This approach addresses the challenges of noisy ASR outputs in critical applications, offering a robust solution to enhance the reliability of clinical dialogue summarization. * Work done † as part of an internship at AICS, § while the authors were at AICS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext baf7add3-5079-4995-bee1-4b70182905ceCited by top-tier papers1
Ask how each one uses itBuilds on7
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- DialogLM: Pre-trained Model for Long Dialogue Understanding and SummarizationMing Zhong, Yang Liu, Yichong Xu, Chenguang Zhu et al.AAAI 2022 · 150 citations
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 54 citations
- Large Language Models are Efficient Learners of Noise-Robust Speech RecognitionYuchen Hu, Chen Chen, Chao-Han Huck Yang, Ruizhe Li et al.ICLR 2024 · 41 citations
Related papers
- Expert-guided Clinical Text Augmentation via Query-Based Model CollaborationDongkyu Cho, Miao Zhang, Gregory Lyng, Rumi ChunaraICML 2026
- PairAug: What Can Augmented Image-Text Pairs Do for Radiology?Yutong Xie, Qi Chen, Sinuo Wang, Minh-Son To et al.CVPR 2024 · 6 citations
- Hearing More with Less: Multi-Modal Retrieval-and-Selection Augmented Conversational LLM-Based ASRBingshen Mu, Hexin Liu, Hongfei Xue, Kun Wei et al.AAAI 2026 · 2 citations
- Interventional Speech Noise Injection for ASR Generalizable Spoken Language UnderstandingYeonJoon Jung, Jaeseong Lee, Seungtaek Choi, Dohyeon Lee et al.EMNLP 2024
- Enhancing Small Medical Learners with Privacy-preserving Contextual PromptingXinlu Zhang, Shiyang Li, Xianjun Yang, Chenxin Tian et al.ICLR 2024 · 13 citations
