Expert-guided Clinical Text Augmentation via Query-Based Model Collaboration
Dongkyu Cho, Miao Zhang, Gregory Lyng, Rumi Chunara
Abstract
Data augmentation is a widely used strategy to improve model robustness and generalization by enriching training datasets with synthetic examples. While large language models (LLMs) have demonstrated strong generative capabilities for this purpose, their applications in high-stakes domains like healthcare present unique challenges due to the risk of generating clinically incorrect or misleading information. In this work, we propose a novel query-based model collaboration framework that integrates expert-level domain knowledge to guide the augmentation process to preserve critical medical information. Compared to existing LLM-based and traditional augmentation methods, our generated data significantly improves preservation of critical medical information and reduces hallucinations at both the token and concept levels. Experiments on downstream clinical prediction tasks demonstrate consistent performance gains over existing augmentation methods. This lightweight collaborative framework addresses the gap between LLM augmentation potential and the safety requirements of specialized domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 197ecc72-0da0-44a0-aec0-da7428d82551Builds on8
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker et al.ICML 2024 · 443 citations
- Improving Commonsense Causal Reasoning by Adversarial Training and Data AugmentationIeva Staliunaite, Philip John Gorinski, Ignacio IacobacciAAAI 2021 · 23 citations
- How Much Data Are Augmentations Worth? An Investigation into Scaling Laws, Invariance, and Implicit RegularizationJonas Geiping, Micah Goldblum, Gowthami Somepalli, Ravid Shwartz-Ziv et al.ICLR 2023 · 11 citations
- HiddenCut: Simple Data Augmentation for Natural Language Understanding with Better GeneralizabilityJiaao Chen, Dinghan Shen, Weizhu Chen, Diyi YangACL 2021
Related papers
- Doctor Approved: Generating Medically Accurate Skin Disease Images through AI-Expert FeedbackJanet Wang, Yunbei Zhang, Zhengming Ding, Jihun HammNeurIPS 2025 · 13 citations
- Enhancing Small Medical Learners with Privacy-preserving Contextual PromptingXinlu Zhang, Shiyang Li, Xianjun Yang, Chenxin Tian et al.ICLR 2024 · 13 citations
- Leveraging Knowledge Graph-Enhanced LLMs for Context-Aware Medical ConsultationSu-Hyeong Park, Ho-Beom Kim, Seong-Jin Park, Dinara Aliyeva et al.EMNLP 2025
- MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language ModelsShrey Pandit, Jiawei Xu, Junyuan Hong, Zhangyang Wang et al.EMNLP 2025 · 8 citations
- A Knowledge-driven Adaptive Collaboration of LLMs for Enhancing Medical Decision-makingXiao Wu, Ting-Zhu Huang, Liang-Jian Deng, Yanyuan Qiao et al.EMNLP 2025 · 1 citation
