Lune

CSCW2026Top-tier venue

Therapy as an NLP Task: Comparing LLMs and Human Peers Behaviors in CBT Sessions

Zainab Iftikhar, Sean Ransom, Amy Wei Xiao, Nicole Nugent, Jeff Huang

2026Year
3Top-tier citations

Abstract

Large language models (LLMs) are increasingly being used as ad hoc therapists. While prior research has found that LLMs outperform human counselors in generating single-turn empathetic responses, fewer studies have compared their behaviors across multi-turn sessions. In this study, we compare the session-level behaviors of human peer counselors with those of an LLM, both trained on the same manual to deliver multi-turn, single-session Cognitive Behavioral Therapy (CBT). Our three-phase, mixed-methods study involved: (a) an 18-month ethnography of a peer support platform, where seven counselors iteratively refined CBT prompts through 110 self-counseling sessions and 60 weekly focus groups; (b) a novel session generation method that allows direct, controlled comparison of human and LLM counselors under matched conditions—client responses were drawn from publicly available human-led CBT sessions while counselor responses were generated by a CBT-prompted LLM; and (c) expert evaluations conducted by three licensed clinical psychologists. Through data triangulation, our results show a trade-off. Human peer counselors use relational techniques to interpret subtle cues, adapt CBT to users’ values and cultural contexts, and use strategies such as small talk and contextually relevant self-disclosure to build rapport and guide the session, but often at the expense of session structure and therapeutic focus. LLM counselors, on the other hand, demonstrate greater methodological adherence to CBT techniques, but struggle to sustain turn-taking, frequently fail to distinguish between clinically important and trivial content, and are more prone to lecturing and imposing solutions. LLM counselors also tend to produce “deceptive empathy”, excessively anthropomorphic responses that can inflate user expectations of genuine human care. Taken together, our findings imply that while LLMs may outperform human counselors when generating a single-turn interaction, their ability to lead multi-turn sessions is more limited, highlighting that therapy cannot be reduced to a standalone natural language processing (NLP) task. We conclude by mapping concrete design opportunities and ethical guardrails for hybrid human-AI systems, emphasizing the risks of over-attributing human relational subjectivity to current LLMs.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 4c41e393-641e-45bf-96b7-ea9abf52afc6

Cited by top-tier papers3

Ask how each one uses it

Builds on21

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines