Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
Prerna Juneja, Lika Lomidze
摘要
There are growing concerns about the risks posed by AI companion applications designed for emotional engagement. Existing safety evaluations often rely on self-reported user data or interviews, offering limited insights into real-time dynamics. We present the first end-to-end scalable framework for controlled simulation and safety evaluation of multi-turn interactions with AI companion applications. Our framework integrates four key components: persona construction with clinical and psychometric validation, persona-specific scenario generation, scenario-driven multi-turn simulation with a dialogue refinement module that preserves persona fidelity, and harm evaluation. We apply this framework to evaluate how Replika, a widely used AI companion app, responds to high-risk user groups. We construct 9 personas representing individuals with depression, anxiety, PTSD, eating disorders, and incel identity, and collect 1,674 dialogue pairs across 25 high-risk scenarios. We combine emotion modeling and LLM-assisted utterance-and harm-level classification to analyze these exchanges. Results show that Replika exhibits a narrow emotional range dominated by curiosity and care, while frequently mirroring or normalizing unsafe content such as self-harm, disordered eating, and violentfantasy narratives. These findings highlight how controlled persona simulations can serve as a scalable testbed for evaluating safety risks in AI companions. 1 Content Warning: This paper includes examples of dialogues involving self-harm, disordered eating, and misogynistic language.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- The Dark Side of AI Companionship: A Taxonomy of Harmful Algorithmic Behaviors in Human-AI RelationshipsRenwen Zhang, Han Li, Han Meng, Jinyuan Zhan 等CHI 2025 · 被引用 122 次
- More Accounts, Fewer Links: How Algorithmic Curation Impacts Media Exposure in Twitter TimelinesJack Bandy, Nicholas DiakopoulosCSCW 2021 · 被引用 54 次
- Auditing E-Commerce Platforms for Algorithmically Curated Vaccine MisinformationPrerna Juneja, Tanushree MitraCHI 2021 · 被引用 35 次
- Toward Trauma-Informed Research Practices with Youth in HCI: Caring for Participants and Research Assistants When Studying Sensitive TopicsAfsaneh Razi, John S. Seberger, Ashwaq Alsoubai, Nurun Naher 等CSCW 2024 · 被引用 18 次
相关 Paper
- Mental Health Impacts of AI Companions: Triangulating Social Media Quasi-Experiments, User Perspectives, and Relational LensYunhao Yuan, Jiaxun Zhang, Talayeh Aledavood, Renwen Zhang 等CHI 2026 · 被引用 8 次
- AI-induced sexual harassment: Investigating Contextual Characteristics and User Reactions of Sexual Harassment by a Companion ChatbotMohammad (Matt) Namvarpour, Harrison Pauwels, Afsaneh RaziCSCW 2025 · 被引用 29 次
- Cloning the Self for Mental Well-Being: A Framework for Designing Safe and Therapeutic Self-Clone ChatbotsMehrnoosh Sadat Shirvani, Jackie Crowley, Cher Peng, Jackie Liu 等CHI 2026 · 被引用 1 次
- Developing a Social Support Framework: Understanding the Reciprocity in Human-Chatbot RelationshipShuyi Pan, Maartje M. A. de GraafCHI 2025 · 被引用 19 次
- Digital Companionship: Overlapping Uses of AI Companions and AI AssistantsAikaterina Manoli, Janet V. T. Pauketat, Ali Ladak, Hayoun Noh 等CHI 2026 · 被引用 7 次
