INTIMA: A Benchmark for Human-AI Companionship Behavior
Lucie-Aimée Kaffee, Giada Pistilli, Yacine Jernite
Abstract
AI companionship, where users develop emotional bonds with AI systems, has emerged as a significant pattern with positive but also concerning implications. We introduce Interactions and Machine Attachment Benchmark (INTIMA), a benchmark for evaluating companionship behaviors in language models. Drawing from psychological theories and user data, we develop a taxonomy of 31 behaviors across four categories and 368 targeted prompts. Responses to these prompts are evaluated as companionship-reinforcing, boundary-maintaining, or neutral. Applying INTIMA to Gemma-3, Phi-4, o4-mini, GPT5-mini, and Claude-4 reveals that companionship-reinforcing behaviors remain much more common across all models, though we observe marked differences between models. Different commercial providers prioritize different categories within the more sensitive parts of the benchmark, which is concerning since both appropriate boundary-setting and emotional support matter for user well-being. These findings highlight the need for more consistent approaches to handling emotionally charged interactions. We release all datasets and evaluation code for our experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on3
- The Dark Side of AI Companionship: A Taxonomy of Harmful Algorithmic Behaviors in Human-AI RelationshipsRenwen Zhang, Han Li, Han Meng, Jinyuan Zhan et al.CHI 2025 · 122 citations
- Large Language Models are not Fair EvaluatorsPeiyi Wang, Lei Li, Liang Chen, Zefan Cai et al.ACL 2024
- DarkBench: Benchmarking Dark Patterns in Large Language ModelsEsben Kran, Jord Nguyen, Akash Kundu, Sami Jawhar et al.ICLR 2025
Related papers
- SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?Yao Dou, Michel Galley, Baolin Peng, Chris Kedzie et al.EMNLP 2025
- Mental Health Impacts of AI Companions: Triangulating Social Media Quasi-Experiments, User Perspectives, and Relational LensYunhao Yuan, Jiaxun Zhang, Talayeh Aledavood, Renwen Zhang et al.CHI 2026 · 8 citations
- Digital Companionship: Overlapping Uses of AI Companions and AI AssistantsAikaterina Manoli, Janet V. T. Pauketat, Ali Ladak, Hayoun Noh et al.CHI 2026 · 7 citations
- Measuring Intent Comprehension in LLMsNadav Kunievsky, James EvansICML 2026 · 1 citation
- Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn ConversationsPrerna Juneja, Lika LomidzeACL 2026
