INTIMA: A Benchmark for Human-AI Companionship Behavior
Lucie-Aimée Kaffee, Giada Pistilli, Yacine Jernite
摘要
AI companionship, where users develop emotional bonds with AI systems, has emerged as a significant pattern with positive but also concerning implications. We introduce Interactions and Machine Attachment Benchmark (INTIMA), a benchmark for evaluating companionship behaviors in language models. Drawing from psychological theories and user data, we develop a taxonomy of 31 behaviors across four categories and 368 targeted prompts. Responses to these prompts are evaluated as companionship-reinforcing, boundary-maintaining, or neutral. Applying INTIMA to Gemma-3, Phi-4, o4-mini, GPT5-mini, and Claude-4 reveals that companionship-reinforcing behaviors remain much more common across all models, though we observe marked differences between models. Different commercial providers prioritize different categories within the more sensitive parts of the benchmark, which is concerning since both appropriate boundary-setting and emotional support matter for user well-being. These findings highlight the need for more consistent approaches to handling emotionally charged interactions. We release all datasets and evaluation code for our experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- The Dark Side of AI Companionship: A Taxonomy of Harmful Algorithmic Behaviors in Human-AI RelationshipsRenwen Zhang, Han Li, Han Meng, Jinyuan Zhan 等CHI 2025 · 被引用 122 次
- Large Language Models are not Fair EvaluatorsPeiyi Wang, Lei Li, Liang Chen, Zefan Cai 等ACL 2024
- DarkBench: Benchmarking Dark Patterns in Large Language ModelsEsben Kran, Jord Nguyen, Akash Kundu, Sami Jawhar 等ICLR 2025
相关 Paper
- SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?Yao Dou, Michel Galley, Baolin Peng, Chris Kedzie 等EMNLP 2025
- Mental Health Impacts of AI Companions: Triangulating Social Media Quasi-Experiments, User Perspectives, and Relational LensYunhao Yuan, Jiaxun Zhang, Talayeh Aledavood, Renwen Zhang 等CHI 2026 · 被引用 8 次
- Digital Companionship: Overlapping Uses of AI Companions and AI AssistantsAikaterina Manoli, Janet V. T. Pauketat, Ali Ladak, Hayoun Noh 等CHI 2026 · 被引用 7 次
- Measuring Intent Comprehension in LLMsNadav Kunievsky, James EvansICML 2026 · 被引用 1 次
- Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn ConversationsPrerna Juneja, Lika LomidzeACL 2026
