CAMEC: Complexity-Aware Multi-Expert Collaboration for Reliable Chinese Medical Question Answering
Yukang Wu, Xiyuan Jia, Jiayi Wu, Hongchen Yu, Yuhan Qiu, Guohua Wu
摘要
Large language models (LLMs) are promising for medical question answering (QA) but remain unreliable in Chinese clinical settings due to hallucinations, weak factual grounding, and difficulty handling clinically complex cases. We propose CAMEC (Complexity-Aware Multi-Expert Collaboration), a framework that combines hierarchical medical adaptation with complexity-aware expert routing for reliable Chinese medical QA. We adopt a three-stage LoRA-based supervised fine-tuning pipeline for domain adaptation, instruction following, and clinical reasoning. At inference, CAMEC routes each query by predicted complexity and selectively recruits three experts: an internal chain-of-thought (CoT) expert, a retrieval-augmented expert over a dense medical vector database, and a knowledge graph (KG) expert over a structured medical knowledge base. An LLM-as-a-Judge module evaluates and critiques expert reports, iteratively refining them into a consensus answer. Experiments on four Chinese medical benchmarks show that CAMEC consistently outperforms strong general and medical LLM baselines, achieving 78.86% (CMExam), 84.15% (MedQA-CN), 78.51% (CMMLU-Med), and 74.40% (CMB-exam), with consistent absolute improvements over the previous state-of-the-art HuatuoGPT-o1-7B across all benchmarks. The complexity-aware router reduces expert invocations and inference cost, making CAMEC both highly effective and computationally efficient.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
- Zhongjing: Enhancing the Chinese Medical Capabilities of Large Language Model through Expert Feedback and Real-World Multi-Turn DialogueSonghua Yang, Hanjie Zhao, Senbin Zhu, Guangyu Zhou 等AAAI 2024 · 被引用 227 次
- MedRAG: Enhancing Retrieval-augmented Generation with Knowledge Graph-Elicited Reasoning for Healthcare CopilotXuejiao Zhao, Siyan Liu, Su-Yin Yang, Chunyan MiaoWWW 2025 · 被引用 134 次
- MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language ModelsYan Cai, Linlin Wang, Ye Wang, Gerard de Melo 等AAAI 2024 · 被引用 42 次
相关 Paper
- Expert-Guided Prompting and Retrieval-Augmented Generation for Emergency Medical Service Question AnsweringXueren Ge, Sahil Murtaza, Anthony Cortez, Homa AlemzadehAAAI 2026 · 被引用 2 次
- CMedCalc-Bench: A Fine-Grained Benchmark for Chinese Medical Calculations in LLMYunyan Zhang, Zhihong Zhu, Xian WuEMNLP 2025 · 被引用 1 次
- From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAGWenhao Wu, Zhentao Tang, Yafu Li, Shixiong Kai 等ICML 2026
- Med-R2: Crafting Trustworthy LLM Physicians via Retrieval and Reasoning of Evidence-Based MedicineKeer Lu, Zheng Liang, Da Pan, Shusen Zhang 等WWW 2026 · 被引用 5 次
- MedLA: A Logic-Driven Multi-Agent Framework for Complex Medical Reasoning with Large Language ModelsSiqi Ma, Jiajie Huang, Fan Zhang, Jinlin Wu 等AAAI 2026 · 被引用 10 次
