Therapy as an NLP Task: Comparing LLMs and Human Peers Behaviors in CBT Sessions
Zainab Iftikhar, Sean Ransom, Amy Wei Xiao, Nicole Nugent, Jeff Huang
Abstract
Large language models (LLMs) are increasingly being used as ad hoc therapists. While prior research has found that LLMs outperform human counselors in generating single-turn empathetic responses, fewer studies have compared their behaviors across multi-turn sessions. In this study, we compare the session-level behaviors of human peer counselors with those of an LLM, both trained on the same manual to deliver multi-turn, single-session Cognitive Behavioral Therapy (CBT). Our three-phase, mixed-methods study involved: (a) an 18-month ethnography of a peer support platform, where seven counselors iteratively refined CBT prompts through 110 self-counseling sessions and 60 weekly focus groups; (b) a novel session generation method that allows direct, controlled comparison of human and LLM counselors under matched conditions—client responses were drawn from publicly available human-led CBT sessions while counselor responses were generated by a CBT-prompted LLM; and (c) expert evaluations conducted by three licensed clinical psychologists. Through data triangulation, our results show a trade-off. Human peer counselors use relational techniques to interpret subtle cues, adapt CBT to users’ values and cultural contexts, and use strategies such as small talk and contextually relevant self-disclosure to build rapport and guide the session, but often at the expense of session structure and therapeutic focus. LLM counselors, on the other hand, demonstrate greater methodological adherence to CBT techniques, but struggle to sustain turn-taking, frequently fail to distinguish between clinically important and trivial content, and are more prone to lecturing and imposing solutions. LLM counselors also tend to produce “deceptive empathy”, excessively anthropomorphic responses that can inflate user expectations of genuine human care. Taken together, our findings imply that while LLMs may outperform human counselors when generating a single-turn interaction, their ability to lead multi-turn sessions is more limited, highlighting that therapy cannot be reduced to a standalone natural language processing (NLP) task. We conclude by mapping concrete design opportunities and ethical guardrails for hybrid human-AI systems, emphasizing the risks of over-attributing human relational subjectivity to current LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4c41e393-641e-45bf-96b7-ea9abf52afc6Cited by top-tier papers3
- Understanding Attitudes and Trust of Generative AI Chatbots for Social Anxiety SupportYimeng Wang, Yinzhou Wang, Kelly Crace, Yixuan ZhangCHI 2025 · 25 citations
- Script-Strategy Aligned Generation: Aligning LLMs with Expert-Crafted Dialogue Scripts and Therapeutic Strategies for PsychotherapyXin Sun, Jan de Wit, Zhuying Li, Jiahuan Pei et al.CSCW 2025 · 6 citations
- Ψ-Arena: Interactive Assessment and Optimization of LLM-based Psychological Counselors with Tripartite FeedbackShijing Zhu, Zhuang Chen, Guanqun Bi, Binghang Li et al.AAAI 2026
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud et al.ICLR 2024 · 762 citations
- "I Hear You, I Feel You": Encouraging Deep Self-disclosure through a ChatbotYi-Chieh Lee, Naomi Yamashita, Yun Huang, Wai FuCHI 2020 · 333 citations
- Mental-LLM: Leveraging Large Language Models for Mental Health Prediction via Online Text DataXuhai Xu, Bingsheng Yao, Yuanzhe Dong, Saadia Gabriel et al.UbiComp 2024 · 281 citations
- Designing a Chatbot as a Mediator for Promoting Deep Self-Disclosure to a Real Mental Health ProfessionalYi-Chieh Lee, Naomi Yamashita, Yun HuangCSCW 2020 · 190 citations
Related papers
- Multi-Level Feedback Generation with Large Language Models for Empowering Novice Peer CounselorsAlicja Chaszczewicz, Raj Sanjay Shah, Ryan Louie, Bruce A. Arnow et al.ACL 2024 · 12 citations
- A Conditional Companion: Lived Experiences of People with Mental Health Disorders Using LLMs: Conditional Companion: LLMs & Mental HealthAditya Kumar Purohit, Hendrik HeuerCHI 2026 · 3 citations
- "Is This Really a Human Peer Supporter?": Misalignments Between Peer Supporters and Experts in LLM-Supported InteractionsKellie Yu Hui Sim, Roy Ka-Wei Lee, Kenny Tsu Wei ChooCSCW 2026
- Large Language Models in Peer-Run Community Behavioral Health Services: Understanding Peer Specialists and Service Users' Perspectives on Opportunities, Risks, and Mitigation StrategiesCindy Peng, Megan Chai, Gao Mo, Naveen Raman et al.CHI 2026 · 2 citations
- What Makes Digital Support Effective? How Therapeutic Skills Affect Clinical Well-BeingWenjie Yang, Anna Fang, Raj Sanjay Shah, Yash Mathur et al.CSCW 2024 · 10 citations
