Can Continual Pretraining Bridge the Performance Gap between General-purpose and Specialized Language Models in the Medical Domain?
Niclas Doll, Jasper Schulze Buschhoff, Shalaka Satheesh, Hammam Abdelwahab, Héctor Allende-Cid, Katrin Klug
摘要
This paper narrows the performance gap between small, specialized models and significantly larger general-purpose models through domain adaptation via continual pre-training and merging. We address the scarcity of specialized non-English data by constructing a high-quality German medical corpus (FineMed-de) from FineWeb2. This corpus is used to continually pre-train and merge three well-known LLMs (ranging from to parameters), creating the DeFineMed model family. A comprehensive evaluation confirms that specialization dramatically enhances model performance on German medical benchmarks. Furthermore, the pairwise win-rate analysis of the Qwen2.5-based models demonstrates an approximately -fold increase in the win-rate against the much larger Mistral-Small-24B-Instruct through domain adaptation. This evidence positions specialized models as a competitive, resource-efficient solution for complex medical instruction-following tasks. While model merging successfully restores instruction-following abilities, a subsequent failure mode analysis reveals inherent trade-offs, including the introduction of language mixing and increased verbosity, highlighting the need for more targeted fine-tuning in future work. This research provides a robust, compliant methodology for developing specialized LLMs, serving as the foundation for practical use in German-speaking healthcare contexts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao 等NeurIPS 2023 · 被引用 475 次
- A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific DiscoveryYu Zhang, Xiusi Chen, Bowen Jin, Sheng Wang 等EMNLP 2024 · 被引用 28 次
- Fewer Truncations Improve Language ModelingHantian Ding, Zijian Wang, Giovanni Paolini, Varun Kumar 等ICML 2024 · 被引用 28 次
相关 Paper
- Efficient Domain Continual pretraining by Mitigating the Stability GapYiduo Guo, Jie Fu, Huishuai Zhang, Dongyan ZhaoACL 2025
- Learning Faster with Better Tokens: Parameter-Efficient Vocabulary Adaptation for Specialized Text SummarizationGunjan Balde, Soumyadeep Roy, Mainack Mondal, Niloy GangulyACL 2026
- DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domainsYanis Labrak, Adrien Bazoge, Richard Dufour, Mickael Rouvier 等ACL 2023 · 被引用 19 次
- Mix-CPT: A Domain Adaptation Framework via Decoupling Knowledge Learning and Format AlignmentJinhao Jiang, Junyi Li, Xin Zhao, Yang Song 等ICLR 2025
- SaulLM-54B & SaulLM-141B: Scaling Up Domain Adaptation for the Legal DomainPierre Colombo, Telmo Pessoa Pires, Malik Boudiaf, Rui Melo 等NeurIPS 2024 · 被引用 58 次
