Bias Similarity Measurement: A Black-Box Audit of Fairness Across LLMs
Hyejun Jeong, Shiqing Ma, Amir Houmansadr
Abstract
Large Language Models (LLMs) reproduce social biases, yet prevailing evaluations score models in isolation, obscuring how biases persist across families and releases. We introduce Bias Similarity Measurement (BSM), which treats fairness as a relational property between models, unifying scalar, distributional, behavioral, and representational signals into a single similarity space. Evaluating 30 LLMs on 1M+ prompts, we find that instruction tuning primarily enforces abstention rather than altering internal representations; small models gain little accuracy and can become less fair under forced choice; and in our evaluation setting, open-weight models can match or exceed proprietary systems. Family signatures diverge: Gemma favors refusal, LLaMA 3.1 approaches neutrality with fewer refusals, and converges toward abstention-heavy behavior overall. Counterintuitively, Gemma 3 Instruct matches GPT-4--level fairness at far lower cost, whereas Gemini’s heavy abstention suppresses utility. Beyond these findings, BSM offers an auditing workflow for procurement, regression testing, and lineage screening, and extends naturally to code and multilingual settings. Our results reframe fairness not as isolated scores but as comparative bias similarity, enabling systematic auditing of LLM ecosystems. Code is available at https://github.com/HyejunJeong/bias_llm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd7a8776-00f3-4b73-97cf-ab69408bf188Builds on11
- Large Language Models are Geographically BiasedRohin Manvi, Samar Khanna, Marshall Burke, David B. Lobell et al.ICML 2024 · 107 citations
- Are You Stealing My Model? Sample Correlation for Fingerprinting Deep Neural NetworksJiyang Guan, Jian Liang, Ran HeNeurIPS 2022 · 57 citations
- ModelDiff: testing-based DNN similarity comparison for model reuse detectionYuanchun Li, Ziqi Zhang, Bingyan Liu, Ziyue Yang et al.ISSTA 2021 · 44 citations
- CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language ModelsNikita Nangia, Clara Vania, Rasika Bhalerao, Samuel R. BowmanEMNLP 2020 · 19 citations
- Similarity Analysis of Contextual Word Representation ModelsJohn M. Wu, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani et al.ACL 2020 · 3 citations
Related papers
- To Mask or to Mirror: Human-AI Alignment in Collective ReasoningCrystal Qian, Aaron T. Parisi, Clémentine Bouleau, Vivian Tsai et al.EMNLP 2025 · 1 citation
- Quantifying Metric and Model Agreement in Bias Evaluation of Large Language ModelsArash Asgari, Huan Wu, Amirreza Naziri, Mojtaba Kolahdouzi et al.ACL 2026
- Reward Models Inherit Value Biases from PretrainingBrian R. Christian, Jessica A. F. Thompson, Elle Michelle Yang, Vincent Adam et al.ICLR 2026 · 4 citations
- Digital Skin, Digital Bias: Uncovering Tone-Based Biases in LLMs and Emoji EmbeddingsMingchen Li, Wajdi Aljedaani, Yingjie Liu, Navyasri Meka et al.WWW 2026
- Pride and Prejudice: LLM Amplifies Self-Bias in Self-RefinementWenda Xu, Guanglei Zhu, Xuandong Zhao, Liangming Pan et al.ACL 2024
