CAMEC: Complexity-Aware Multi-Expert Collaboration for Reliable Chinese Medical Question Answering
Yukang Wu, Xiyuan Jia, Jiayi Wu, Hongchen Yu, Yuhan Qiu, Guohua Wu
Abstract
Large language models (LLMs) are promising for medical question answering (QA) but remain unreliable in Chinese clinical settings due to hallucinations, weak factual grounding, and difficulty handling clinically complex cases. We propose CAMEC (Complexity-Aware Multi-Expert Collaboration), a framework that combines hierarchical medical adaptation with complexity-aware expert routing for reliable Chinese medical QA. We adopt a three-stage LoRA-based supervised fine-tuning pipeline for domain adaptation, instruction following, and clinical reasoning. At inference, CAMEC routes each query by predicted complexity and selectively recruits three experts: an internal chain-of-thought (CoT) expert, a retrieval-augmented expert over a dense medical vector database, and a knowledge graph (KG) expert over a structured medical knowledge base. An LLM-as-a-Judge module evaluates and critiques expert reports, iteratively refining them into a consensus answer. Experiments on four Chinese medical benchmarks show that CAMEC consistently outperforms strong general and medical LLM baselines, achieving 78.86% (CMExam), 84.15% (MedQA-CN), 78.51% (CMMLU-Med), and 74.40% (CMB-exam), with consistent absolute improvements over the previous state-of-the-art HuatuoGPT-o1-7B across all benchmarks. The complexity-aware router reduces expert invocations and inference cost, making CAMEC both highly effective and computationally efficient.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fbd65cb7-e306-4540-8b8b-ba215f1fa032Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Zhongjing: Enhancing the Chinese Medical Capabilities of Large Language Model through Expert Feedback and Real-World Multi-Turn DialogueSonghua Yang, Hanjie Zhao, Senbin Zhu, Guangyu Zhou et al.AAAI 2024 · 227 citations
- MedRAG: Enhancing Retrieval-augmented Generation with Knowledge Graph-Elicited Reasoning for Healthcare CopilotXuejiao Zhao, Siyan Liu, Su-Yin Yang, Chunyan MiaoWWW 2025 · 134 citations
- MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language ModelsYan Cai, Linlin Wang, Ye Wang, Gerard de Melo et al.AAAI 2024 · 42 citations
Related papers
- Expert-Guided Prompting and Retrieval-Augmented Generation for Emergency Medical Service Question AnsweringXueren Ge, Sahil Murtaza, Anthony Cortez, Homa AlemzadehAAAI 2026 · 2 citations
- CMedCalc-Bench: A Fine-Grained Benchmark for Chinese Medical Calculations in LLMYunyan Zhang, Zhihong Zhu, Xian WuEMNLP 2025 · 1 citation
- From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAGWenhao Wu, Zhentao Tang, Yafu Li, Shixiong Kai et al.ICML 2026
- Med-R2: Crafting Trustworthy LLM Physicians via Retrieval and Reasoning of Evidence-Based MedicineKeer Lu, Zheng Liang, Da Pan, Shusen Zhang et al.WWW 2026 · 5 citations
- MedLA: A Logic-Driven Multi-Agent Framework for Complex Medical Reasoning with Large Language ModelsSiqi Ma, Jiajie Huang, Fan Zhang, Jinlin Wu et al.AAAI 2026 · 10 citations
