Towards AI-Assisted Psychotherapy: Emotion-Guided Generative Interventions
Kilichbek Haydarov, Youssef Mohamed, Emilio Goldenhersch, Paul OCallaghan, Li-jia Li, Mohamed Elhoseiny
Abstract
Large language models (LLMs) hold promise for therapeutic interventions, yet most existing datasets rely solely on text, overlooking nonverbal emotional cues essential to real-world therapy. To address this, we introduce a multimodal dataset of 1,441 publicly sourced therapy session videos containing both dialogue and non-verbal signals such as facial expressions and vocal tone. Inspired by Hochschild's concept of emotional labor, we propose a computational formulation of emotional dissonance-the mismatch between facial and vocal emotion-and use it to guide emotionally aware prompting. Our experiments show that integrating multimodal cues, especially dissonance, improves the quality of generated interventions. We also find that LLM-based evaluators misalign with expert assessments in this domain, highlighting the need for humancentered evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui et al.EMNLP 2024 · 231 citations
- Towards Interpretable Mental Health Analysis with Large Language ModelsKailai Yang, Shaoxiong Ji, Tianlin Zhang, Qianqian Xie et al.EMNLP 2023 · 114 citations
Related papers
- MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with ResistanceSubin Kim, Hoonrae Kim, Jihyun Lee, Yejin Jeon et al.EMNLP 2025 · 6 citations
- Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face ConversationSe Jin Park, Chae Won Kim, Hyeongseop Rha, Minsu Kim et al.ACL 2024
- Emotion-Coherent Reasoning for Multimodal LLMs via Emotional Rationale VerifierHyeongseop Rha, Jeong Hun Yeo, Yeonju Kim, Yong Man RoAAAI 2026 · 1 citation
- Towards Aligning Multimodal LLMs with Human Experts: A Focus on Parent-Child InteractionWeiyan Shi, Kenny Tsu Wei ChooCHI 2026 · 2 citations
- An LLM-based Simulation Framework for Embodied Conversational Agents in Psychological CounselingLixiu Wu, Yuanrong Tang, Qisen Pan, Xianyang Zhan et al.AAAI 2026 · 1 citation
