Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
Prerna Juneja, Lika Lomidze
Abstract
There are growing concerns about the risks posed by AI companion applications designed for emotional engagement. Existing safety evaluations often rely on self-reported user data or interviews, offering limited insights into real-time dynamics. We present the first end-to-end scalable framework for controlled simulation and safety evaluation of multi-turn interactions with AI companion applications. Our framework integrates four key components: persona construction with clinical and psychometric validation, persona-specific scenario generation, scenario-driven multi-turn simulation with a dialogue refinement module that preserves persona fidelity, and harm evaluation. We apply this framework to evaluate how Replika, a widely used AI companion app, responds to high-risk user groups. We construct 9 personas representing individuals with depression, anxiety, PTSD, eating disorders, and incel identity, and collect 1,674 dialogue pairs across 25 high-risk scenarios. We combine emotion modeling and LLM-assisted utterance-and harm-level classification to analyze these exchanges. Results show that Replika exhibits a narrow emotional range dominated by curiosity and care, while frequently mirroring or normalizing unsafe content such as self-harm, disordered eating, and violentfantasy narratives. These findings highlight how controlled persona simulations can serve as a scalable testbed for evaluating safety risks in AI companions. 1 Content Warning: This paper includes examples of dialogues involving self-harm, disordered eating, and misogynistic language.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d583bb1-a145-4dd4-9b9d-afb2b47cd717Builds on8
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- The Dark Side of AI Companionship: A Taxonomy of Harmful Algorithmic Behaviors in Human-AI RelationshipsRenwen Zhang, Han Li, Han Meng, Jinyuan Zhan et al.CHI 2025 · 122 citations
- More Accounts, Fewer Links: How Algorithmic Curation Impacts Media Exposure in Twitter TimelinesJack Bandy, Nicholas DiakopoulosCSCW 2021 · 54 citations
- Auditing E-Commerce Platforms for Algorithmically Curated Vaccine MisinformationPrerna Juneja, Tanushree MitraCHI 2021 · 35 citations
- Toward Trauma-Informed Research Practices with Youth in HCI: Caring for Participants and Research Assistants When Studying Sensitive TopicsAfsaneh Razi, John S. Seberger, Ashwaq Alsoubai, Nurun Naher et al.CSCW 2024 · 18 citations
Related papers
- Mental Health Impacts of AI Companions: Triangulating Social Media Quasi-Experiments, User Perspectives, and Relational LensYunhao Yuan, Jiaxun Zhang, Talayeh Aledavood, Renwen Zhang et al.CHI 2026 · 8 citations
- AI-induced sexual harassment: Investigating Contextual Characteristics and User Reactions of Sexual Harassment by a Companion ChatbotMohammad (Matt) Namvarpour, Harrison Pauwels, Afsaneh RaziCSCW 2025 · 29 citations
- Cloning the Self for Mental Well-Being: A Framework for Designing Safe and Therapeutic Self-Clone ChatbotsMehrnoosh Sadat Shirvani, Jackie Crowley, Cher Peng, Jackie Liu et al.CHI 2026 · 1 citation
- Developing a Social Support Framework: Understanding the Reciprocity in Human-Chatbot RelationshipShuyi Pan, Maartje M. A. de GraafCHI 2025 · 19 citations
- Digital Companionship: Overlapping Uses of AI Companions and AI AssistantsAikaterina Manoli, Janet V. T. Pauketat, Ali Ladak, Hayoun Noh et al.CHI 2026 · 7 citations
