User-Driven Value Alignment: Understanding Users' Perceptions and Strategies for Addressing Biased and Discriminatory Statements in AI Companions
Xianzhe Fan, Qing Xiao, Xuhui Zhou, Jiaxin Pei, Maarten Sap, Zhicong Lu, Hong Shen
摘要
Content Warning: This paper presents textual examples that may be offensive or upsetting.
Large language model-based AI companions are increasingly viewed by users as friends or romantic partners, leading to deep emotional bonds. However, they can generate biased, discriminatory, and harmful outputs. Recently, users are taking the initiative to address these harms and re-align AI companions. We introduce the concept of user-driven value alignment, where users actively identify, challenge, and attempt to correct AI outputs they perceive as harmful, aiming to guide the AI to better align with their values. We analyzed 77 social media posts about discriminatory AI statements and conducted semi-structured interviews with 20 experienced users. Our analysis revealed six common types of discriminatory statements perceived by users, how users make sense of those AI behaviors, and seven user-driven alignment strategies, such as gentle persuasion and anger expression. We discuss implications for supporting user-driven value alignment in future AI systems, where users and their communities have greater agency.
• Human-centered computing → Empirical studies in HCI .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Interaction Context Often Increases Sycophancy in LLMsShomik Jain, Charlotte Park, Matt Viana, Ashia Wilson 等CHI 2026 · 被引用 12 次
- Relational Dissonance in Human-AI Interactions: The Case of Knowledge WorkEmrecan Gulay, Eleonora Picco, Enrico Glerean, Corinna CoupetteCHI 2026 · 被引用 8 次
- Negotiating Digital Identities with AI Companions: Motivations, Strategies, and Emotional OutcomesRenkai Ma, Shuo Niu, Lingyao Li, Alex Hirth 等CHI 2026 · 被引用 7 次
- AI and My Values: User Perceptions of LLMs' Ability to Extract, Embody, and Explain Human Values from Casual ConversationsBhada Yun, Renn Su, April Yi WangCHI 2026 · 被引用 7 次
- POET: Supporting Prompting Creativity and Personalization with Automated Expansion of Text-to-Image GenerationEvans Xu Han, Alice Qian Zhang, Haiyi Zhu, Hong Shen 等UIST 2025 · 被引用 5 次
它引用的顶会 Paper27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer 等NeurIPS 2023 · 被引用 1,486 次
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng 等EMNLP 2024 · 被引用 479 次
- Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human SupervisionZhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang 等NeurIPS 2023 · 被引用 463 次
相关 Paper
- Unintended Harms of Value-Aligned LLMs: Psychological and Empirical InsightsSooyung Choi, Jaehyeok Lee, Xiaoyuan Yi, Jing Yao 等ACL 2025
- The Typing Cure: Experiences with Large Language Model Chatbots for Mental Health SupportInhwa Song, Sachin R. Pendse, Neha Kumar, Munmun De ChoudhuryCSCW 2025 · 被引用 41 次
- Caught in a Mafia Romance: How Users Explore Intimate Narratives with ChatbotsJulia B. Kieserman, Cat Mai, Sara Lignell, Lucy Qin 等CHI 2026 · 被引用 1 次
- Toward User-Driven Algorithm Auditing: Investigating users' strategies for uncovering harmful algorithmic behaviorAlicia DeVos, Aditi Dhabalia, Hong Shen, Kenneth Holstein 等CHI 2022 · 被引用 96 次
- "Please, don't kill the only model that still feels human": Understanding the #Keep4o BacklashHuiqian LaiCHI 2026 · 被引用 5 次
