Certifying Counterfactual Bias in LLMs
Isha Chaudhary, Qian Hu, Manoj Kumar, Morteza Ziyadi, Rahul Gupta, Gagandeep Singh
摘要
Warning: This paper contains model outputs that are offensive in nature. Large Language Models (LLMs) can produce biased responses that can cause representational harms. However, conventional studies are insufficient to thoroughly evaluate biases across LLM responses for different demographic groups (a.k.a. counterfactual bias), as they do not scale to large number of inputs and do not provide guarantees. Therefore, we propose the first framework, LLMCert-B that certifies LLMs for counterfactual bias on distributions of prompts. A certificate consists of high-confidence bounds on the probability of unbiased LLM responses for any set of counterfactual prompts -prompts differing by demographic groups, sampled from a distribution. We illustrate counterfactual bias certification for distributions of counterfactual prompts created by applying prefixes sampled from prefix distributions, to a given set of prompts. We consider prefix distributions consisting random token sequences, mixtures of manual jailbreaks, and perturbations of jailbreaks in LLM's embedding space. We generate non-trivial certificates for SOTA LLMs, exposing their vulnerabilities over distributions of prompts generated from computationally inexpensive prefix distributions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- BiasBusters: Uncovering and Mitigating Tool Selection Bias in Large Language ModelsThierry Blankenstein, Jialin Yu, Zixuan Li, Vassilis Plachouras 等ICLR 2026 · 被引用 8 次
- Inertia in Moral and Value Judgments of Large Language ModelsBruce W. Lee, Yeongheon Lee, Hyunsoo ChoACL 2026 · 被引用 5 次
- Bias Similarity Measurement: A Black-Box Audit of Fairness Across LLMsHyejun Jeong, Shiqing Ma, Amir HoumansadrICLR 2026 · 被引用 1 次
- Investigating Counterfactual Unfairness in LLMs towards Identities through HumorShubin Kim, Yejin Son, Junyeong Park, Keummin Ka 等ACL 2026
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- Beta-CROWN: Efficient Bound Propagation with Per-neuron Split Constraints for Neural Network Robustness VerificationShiqi Wang, Huan Zhang, Kaidi Xu, Xue Lin 等NeurIPS 2021 · 被引用 359 次
- Conformal Language ModelingVictor Quach, Adam Fisch, Tal Schuster, Adam Yala 等ICLR 2024 · 被引用 132 次
相关 Paper
- Adaptive Generation of Bias-Eliciting Questions for LLMsRobin Staab, Jasper Dekoninck, Maximilian Baader, Martin VechevICML 2026
- How Catastrophic is Your LLM? Certifying Risks in ConversationChengxiao Wang, Isha Chaudhary, Qian Hu, Weitong Ruan 等ICLR 2026 · 被引用 1 次
- Bias Association Discovery Framework for Open-Ended LLM GenerationsJinhao Pan, Chahat Raj, Ziwei ZhuAAAI 2026 · 被引用 1 次
- The Impossibility of Fair LLMsJacy Reese Anthis, Kristian Lum, Michael D. Ekstrand, Avi Feller 等ACL 2025
- Shh, don't say that! Domain Certification in LLMsCornelius Emde, Alasdair Paren, Preetham Arvind, Maxime Guillaume Kayser 等ICLR 2025
