BiasAsker: Measuring the Bias in Conversational AI System
Yuxuan Wan, Wenxuan Wang, Pinjia He, Jiazhen Gu, Haonan Bai, Michael R. Lyu
摘要
Powered by advanced Artificial Intelligence (AI) techniques, conversational AI systems, such as ChatGPT and digital assistants like Siri, have been widely deployed in daily life. However, such systems may still produce content containing biases and stereotypes, causing potential social problems. Due to the data-driven, black-box nature of modern AI techniques, comprehensively identifying and measuring biases in conversational systems remains a challenging task. Particularly, it is hard to generate inputs that can comprehensively trigger potential bias due to the lack of data containing both social groups as well as biased properties. In addition, modern conversational systems can produce diverse responses (e.g., chatting and explanation), which makes existing bias detection methods simply based on the sentiment and the toxicity hardly being adopted. In this paper, we propose BiasAsker, an automated framework to identify and measure social bias in conversational AI systems. To obtain social groups and biased properties, we construct a comprehensive social bias dataset, containing a total of 841 groups and 8,110 biased properties. Given the dataset, BiasAsker automatically generates questions and adopts a novel method based on existence measurement to identify two types of biases (i.e., absolute bias and related bias) in conversational systems. Extensive experiments on 8 commercial systems and 2 famous research models, such as Chat-GPT and GPT-3, show that 32.83% of the questions generated by BiasAsker can trigger biased behaviors in these widely deployed conversational systems. All the code, data, and experimental results have been released to facilitate future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Measuring Political Bias in Large Language Models: What Is Said and How It Is SaidYejin Bang, Delong Chen, Nayeon Lee, Pascale FungACL 2024 · 被引用 21 次
- VisBias: Measuring Explicit and Implicit Social Biases in Vision Language ModelsJen-Tse Huang, Jiantong Qin, Jianping Zhang, Youliang Yuan 等EMNLP 2025 · 被引用 13 次
- Glitch Tokens in Large Language Models: Categorization Taxonomy and Effective DetectionYuxi Li, Yi Liu, Gelei Deng, Ying Zhang 等FSE 2024 · 被引用 12 次
- An Image is Worth a Thousand Toxic Words: A Metamorphic Testing Framework for Content Moderation SoftwareWenxuan Wang, Jingyuan Huang, Jen-tse Huang, Chang Chen 等ASE 2023 · 被引用 7 次
- VLATest: Testing and Evaluating Vision-Language-Action Models for Robotic ManipulationZhijie Wang, Zhehua Zhou, Jiayang Song, Yuheng Huang 等FSE 2025 · 被引用 6 次
它引用的顶会 Paper17
- Hidden Voice CommandsNicholas Carlini, Pratyush Mishra, Tavish Vaidya, Yuankai Zhang 等USENIX Security 2016 · 被引用 672 次
- Improving Adversarial Transferability via Neuron Attribution-based AttacksJianping Zhang, Weibin Wu, Jen-tse Huang, Yizhan Huang 等CVPR 2022 · 被引用 140 次
- DeepCrime: mutation testing of deep learning systems based on real faultsNargiz Humbatova, Gunel Jahangirova, Paolo TonellaISSTA 2021 · 被引用 114 次
- RobOT: Robustness-Oriented Testing for Deep Learning SystemsJingyi Wang, Jialuo Chen, Youcheng Sun, Xingjun Ma 等ICSE 2021 · 被引用 62 次
- Mitigating Gender Bias for Neural Dialogue Generation with Adversarial LearningHaochen Liu, Wentao Wang, Yiqi Wang, Hui Liu 等EMNLP 2020 · 被引用 55 次
相关 Paper
- Uncovering and Quantifying Social Biases in Code GenerationYan Liu, Xiaokang Chen, Yan Gao, Zhe Su 等NeurIPS 2023 · 被引用 47 次
- RedditBias: A Real-World Resource for Bias Evaluation and Debiasing of Conversational Language ModelsSoumya Barikeri, Anne Lauscher, Ivan Vulic, Goran GlavasACL 2021
- CHBias: Bias Evaluation and Mitigation of Chinese Conversational Language ModelsJiaxu Zhao, Meng Fang, Zijing Shi, Yitong Li 等ACL 2023 · 被引用 11 次
- GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark ConstructionVirginia K. Felkner, Jennifer A. Thompson, Jonathan MayACL 2024 · 被引用 3 次
- "I followed what felt right, not what I was told": Autonomy, Coaching, and Recognizing Bias Through AI-Mediated DialogueAtieh Taheri, Hamza El Alaoui, Patrick Carrington, Jeffrey P. BighamCHI 2026 · 被引用 1 次
