MentalSeek-Dx: Towards Progressive Hypothetico-Deductive Reasoning for Real-world Psychiatric Diagnosis
Xiao Sun, Yuming Yang, Xinyi Jiang, Yu Tian, Junnan Zhu, Jiang Zhong, Qin Lei, Jingwang Huang, Haoyang Zeng, Xinyu Zhou, Xin Xiao, Kaiwen Wei
摘要
Mental health disorders represent a burgeoning global public health challenge. While Large Language Models (LLMs) have demonstrated potential in psychiatric assessment, their clinical utility is severely constrained by benchmarks that lack ecological validity and finegrained diagnostic supervision. To bridge this gap, we introduce MentalDx Bench, the first benchmark dedicated to disorder-level psychiatric diagnosis within real-world clinical settings. Comprising 712 de-identified electronic health records annotated by boardcertified psychiatrists under ICD-11 guidelines, the benchmark covers 76 disorders across 16 diagnostic categories. Evaluation of 18 LLMs reveals a critical paradigm misalignment: strong performance at coarse diagnostic categorization contrasts with systematic failure at disorder-level diagnosis, underscoring a gap between pattern-based modeling and clinical hypothetico-deductive reasoning. In response, we propose MentalSeek-Dx, a medical-specialized LLM trained to internalize this clinical reasoning process through supervised trajectory construction and curriculumbased reinforcement learning. Experiments on MentalDx Bench demonstrate that MentalSeek-Dx achieves state-of-the-art (SOTA) performance with only 14B parameters, establishing a clinically grounded framework for reliable psychiatric diagnosis. The dataset and code are available at MentalSeek-Dx.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- D4: a Chinese Dialogue Dataset for Depression-Diagnosis-Oriented ChatBinwei Yao, Chao Shi, Likai Zou, Lingfeng Dai 等EMNLP 2022 · 被引用 23 次
- Psyche-R1: Towards Reliable Psychological LLMs through Unified Empathy, Expertise, and ReasoningChongyuan Dai, Jinpeng Hu, Hongchang Shi, Zhuo Li 等ACL 2026 · 被引用 18 次
- Still Not Quite There! Evaluating Large Language Models for Comorbid Mental Health DiagnosisAmey Hengle, Atharva Kulkarni, Shantanu Patankar, Madhumitha Chandrasekaran 等EMNLP 2024 · 被引用 1 次
- Competence-based Multimodal Curriculum Learning for Medical Report GenerationFenglin Liu, Shen Ge, Xian WuACL 2021
相关 Paper
- DDxTutor: Clinical Reasoning Tutoring System with Differential Diagnosis-Based Structured ReasoningQian Wu, Zheyao Gao, Longfei Gou, Qi DouACL 2025 · 被引用 2 次
- MDD-5k: A New Diagnostic Conversation Dataset for Mental Disorders Synthesized via Neuro-Symbolic LLM AgentsCongchi Yin, Feng Li, Shu Zhang, Zike Wang 等AAAI 2025 · 被引用 18 次
- Moving Beyond Medical Exams: A Clinician-Annotated Fairness Dataset of Real-World Tasks and Ambiguity in Mental HealthcareMax Lamparth, Declan Grabb, Amy Franks, Scott Gershan 等ICLR 2026 · 被引用 5 次
- CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question AnsweringYahan Li, Jifan Yao, John Bosco S. Bunyi, Adam C. Frank 等ICLR 2026 · 被引用 24 次
- REACT-LLM: A Benchmark for Evaluating LLM Integration with Causal Features in Clinical Prognostic TasksLinna Wang, Zhixuan You, Qihui Zhang, Jiunan Wen 等AAAI 2026
