Using the Crowd to Prevent Harmful AI Behavior
Travis Mandel, Jahnu Best, Randall H. Tanaka, Hiram Temple, Chansen Haili, Sebastian J. Carter, Kayla Schlechtinger, Roy Szeto
摘要
To prevent harmful AI behavior, people need to specify constraints that forbid undesirable actions. Unfortunately, this is a complex task, since writing rules that distinguish harmful from non-harmful actions tends to be quite difficult in real-world situations. Therefore, such decisions have historically been made by a small group of powerful AI companies and developers, with limited community input. In this paper, we study how to enable a crowd of non-AI experts to work together to communicate high-quality, reliable constraints to AI systems. We first focus on understanding how humans reason about temporal dynamics in the context of AI behavior, finding through experiments on a novel game-based testbed that participants tend to adopt a long-term notion of harm, even in uncertain situations that do not affect them directly. Building off of this insight, we explore task design for long-term constraint specification, developing new filtering approaches and new methods of promoting user reflection. Next, we develop a novel rule-based interface which allows people to craft rules in an accessible fashion without programming knowledge. We test our approaches on a real-world AI problem in the domain of education, and find that our new filtering mechanisms and interfaces significantly improve constraint quality and human efficiency. We also demonstrate how these systems can be applied to other real-world AI problems (e.g. in social networks).
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- Towards Transparency in Dermatology Image Datasets with Skin Tone Annotations by Experts, Crowds, and an AlgorithmMatthew Groh, Caleb Harris, Roxana Daneshjou, Omar Badri 等CSCW 2022 · 被引用 62 次
- Discovering and Validating AI Errors With Crowdsourced Failure ReportsÁngel Alexander Cabrera, Abraham J. Druck, Jason I. Hong, Adam PererCSCW 2021 · 被引用 60 次
- Power and Play: Investigating "License to Critique" in Teams' AI Ethics DiscussionsDavid Gray Widder, Laura Dabbish, James D. Herbsleb, Nikolas MartelaroCSCW 2024 · 被引用 13 次
相关 Paper
- AI Knowledge: Improving AI Delegation through Human EnablementMarc Pinski, Martin Adam, Alexander BenlianCHI 2023 · 被引用 58 次
- Worker Discretion Advised: Co-designing Risk Disclosure in Crowdsourced Responsible AI (RAI) Content WorkAlice Qian, Ziqi Yang, Ryland Shaw, Jina Suh 等CHI 2026 · 被引用 1 次
- Locating Risk: Task Designers and the Challenge of Risk Disclosure in Crowdsourced RAI Content Work CSCW029Alice Qian, Ryland Shaw, Laura Dabbish, Jina Suh 等CSCW 2026
- Multi-Round Human–AI Collaboration with User-Specified RequirementsSima Noorani, Shayan Kiyani, Hamed Hassani, George PappasICML 2026 · 被引用 2 次
- Teaching Humans When to Defer to a Classifier via ExemplarsHussein Mozannar, Arvind Satyanarayan, David A. SontagAAAI 2022 · 被引用 49 次
