Deciphering Cognitive Distortions in Patient-Doctor Mental Health Conversations: A Multimodal LLM-Based Detection and Reasoning Framework
Gopendra Vikram Singh, Sai Vemulapalli, Mauajama Firdaus, Asif Ekbal
Abstract
Cognitive distortion research holds increasing significance as it sheds light on pervasive errors in thinking patterns, providing crucial insights into mental health challenges and fostering the development of targeted interventions and therapies. This paper delves into the complex domain of cognitive distortions which are prevalent distortions in cognitive processes often associated with mental health issues. Focusing on patient-doctor dialogues, we introduce a pioneering method for detecting and reasoning about cognitive distortions utilizing Large Language Models (LLMs). Operating within a multimodal context encompassing audio, video, and textual data, our approach underscores the critical importance of integrating diverse modalities for a comprehensive understanding of cognitive distortions. By leveraging multimodal information, including audio, video, and textual data, our method offers a nuanced perspective that enhances the accuracy and depth of cognitive distortion detection and reasoning in a zero-shot manner. Our proposed hierarchical framework adeptly tackles both detection and reasoning tasks, showcasing significant performance enhancements compared to current methodologies. Through comprehensive analysis, we elucidate the efficacy of our approach, offering promising insights into the diagnosis and understanding of cognitive distortions in multimodal settings.The code and dataset can be found here: https://www.iitp.ac. in/ ai-nlp-ml/resources.html#ZS-CoDR . CoD Reasoning: The patient's final words, "They are always commenting on everything that I'm doing." could be seen as an example of cognitive distortion. This distortion occurs when someone assumes others are always scrutinizing them, despite lacking evidence. The patient's belief that others are constantly monitoring and critiquing their actions is exaggerated and unsupported, demonstrating a distorted perception of external attention. D: Ok, ok. And can you hear what they are actually saying? P: Yeah, they are talking about me.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- DDO: Dual-Decision Optimization for LLM-Based Medical Consultation via Multi-Agent CollaborationZhihao Jia, Mingyi Jia, Junwen Duan, Jian-xin WangEMNLP 2025 · 2 citations
- Towards AI-Assisted Psychotherapy: Emotion-Guided Generative InterventionsKilichbek Haydarov, Youssef Mohamed, Emilio Goldenhersch, Paul OCallaghan et al.EMNLP 2025
Builds on9
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- EMOCA: Emotion Driven Monocular Face Capture and AnimationRadek Danecek, Michael J. Black, Timo BolkartCVPR 2022 · 180 citations
- Explaining (Sarcastic) Utterances to Enhance Affect Understanding in Multimodal DialoguesShivani Kumar, Ishani Mondal, Md. Shad Akhtar, Tanmoy ChakrabortyAAAI 2023 · 22 citations
Related papers
- Multi-View Attention Multiple-Instance Learning Enhanced by LLM Reasoning for Cognitive Distortion DetectionJun Seo Kim, Hyemi Kim, Woo Joo Oh, Hongjin Cho et al.ACL 2026
- Cognitive Policy-Driven LLM for Diagnosis and Intervention of Cognitive Distortions in Emotional Support ConversationLin Zhong, Renjin Zhu, Shujuan Ma, Jinhao Cui et al.ACL 2026
- From Classification to Clinical Insights: Towards Analyzing and Reasoning About Mobile and Behavioral Health Data With Large Language ModelsZachary Englhardt, Chengqian Ma, Margaret E. Morris, Chun-Cheng Chang et al.UbiComp 2024 · 51 citations
- MIND: Empowering Mental Health Clinicians with Multimodal Data Insights through a Narrative DashboardRuishi Zou, Shiyu Xu, Margaret E. Morris, Jihan Ryu et al.CHI 2026 · 2 citations
- When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression DetectionXiangyu Zhang, Hexin Liu, Kaishuai Xu, Qiquan Zhang et al.EMNLP 2024 · 13 citations
