MentalSeek-Dx: Towards Progressive Hypothetico-Deductive Reasoning for Real-world Psychiatric Diagnosis
Xiao Sun, Yuming Yang, Xinyi Jiang, Yu Tian, Junnan Zhu, Jiang Zhong, Qin Lei, Jingwang Huang, Haoyang Zeng, Xinyu Zhou, Xin Xiao, Kaiwen Wei
Abstract
Mental health disorders represent a burgeoning global public health challenge. While Large Language Models (LLMs) have demonstrated potential in psychiatric assessment, their clinical utility is severely constrained by benchmarks that lack ecological validity and finegrained diagnostic supervision. To bridge this gap, we introduce MentalDx Bench, the first benchmark dedicated to disorder-level psychiatric diagnosis within real-world clinical settings. Comprising 712 de-identified electronic health records annotated by boardcertified psychiatrists under ICD-11 guidelines, the benchmark covers 76 disorders across 16 diagnostic categories. Evaluation of 18 LLMs reveals a critical paradigm misalignment: strong performance at coarse diagnostic categorization contrasts with systematic failure at disorder-level diagnosis, underscoring a gap between pattern-based modeling and clinical hypothetico-deductive reasoning. In response, we propose MentalSeek-Dx, a medical-specialized LLM trained to internalize this clinical reasoning process through supervised trajectory construction and curriculumbased reinforcement learning. Experiments on MentalDx Bench demonstrate that MentalSeek-Dx achieves state-of-the-art (SOTA) performance with only 14B parameters, establishing a clinically grounded framework for reliable psychiatric diagnosis. The dataset and code are available at MentalSeek-Dx.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ae6fae5-9b33-465a-9cbf-8c4aab9f5b22Builds on4
- D4: a Chinese Dialogue Dataset for Depression-Diagnosis-Oriented ChatBinwei Yao, Chao Shi, Likai Zou, Lingfeng Dai et al.EMNLP 2022 · 23 citations
- Psyche-R1: Towards Reliable Psychological LLMs through Unified Empathy, Expertise, and ReasoningChongyuan Dai, Jinpeng Hu, Hongchang Shi, Zhuo Li et al.ACL 2026 · 18 citations
- Still Not Quite There! Evaluating Large Language Models for Comorbid Mental Health DiagnosisAmey Hengle, Atharva Kulkarni, Shantanu Patankar, Madhumitha Chandrasekaran et al.EMNLP 2024 · 1 citation
- Competence-based Multimodal Curriculum Learning for Medical Report GenerationFenglin Liu, Shen Ge, Xian WuACL 2021
Related papers
- DDxTutor: Clinical Reasoning Tutoring System with Differential Diagnosis-Based Structured ReasoningQian Wu, Zheyao Gao, Longfei Gou, Qi DouACL 2025 · 2 citations
- MDD-5k: A New Diagnostic Conversation Dataset for Mental Disorders Synthesized via Neuro-Symbolic LLM AgentsCongchi Yin, Feng Li, Shu Zhang, Zike Wang et al.AAAI 2025 · 18 citations
- Moving Beyond Medical Exams: A Clinician-Annotated Fairness Dataset of Real-World Tasks and Ambiguity in Mental HealthcareMax Lamparth, Declan Grabb, Amy Franks, Scott Gershan et al.ICLR 2026 · 5 citations
- CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question AnsweringYahan Li, Jifan Yao, John Bosco S. Bunyi, Adam C. Frank et al.ICLR 2026 · 24 citations
- REACT-LLM: A Benchmark for Evaluating LLM Integration with Causal Features in Clinical Prognostic TasksLinna Wang, Zhixuan You, Qihui Zhang, Jiunan Wen et al.AAAI 2026
