STILE: Exploring and Debugging Social Biases in Pre-trained Text Representations
Samia Kabir, Lixiang Li, Tianyi Zhang
Abstract
The recent success of Natural Language Processing (NLP) relies heavily on pre-trained text representations such as word embeddings. However, pre-trained text representations may exhibit social biases and stereotypes, e.g., disproportionately associating gender with occupations. Though prior work presented various bias detection algorithms, they are limited to pre-defined biases and lack effective interaction support. In this work, we propose Stile, an interactive system that supports mixed-initiative bias discovery and debugging in pre-trained text representations. Stile provides users the flexibility to interactively define and customize biases to detect based on their interests. Furthermore, it provides a bird’s-eye view of detected biases in a Chord diagram and allows users to dive into the training data to investigate how a bias was developed. Our lab study and expert review confirm the usefulness and usability of Stile as an effective aid in identifying and understanding biases in pre-trained text representations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d14c77c8-be1a-438d-97da-3faf4269bc66Cited by top-tier papers4
- User-Driven Value Alignment: Understanding Users' Perceptions and Strategies for Addressing Biased and Discriminatory Statements in AI CompanionsXianzhe Fan, Qing Xiao, Xuhui Zhou, Jiaxin Pei et al.CHI 2025 · 15 citations
- Investigating the Capabilities and Limitations of Machine Learning for Identifying Bias in English Language Data with Information and Heritage ProfessionalsLucy Havens, Benjamin Bach, Melissa Terras, Beatrice AlexCHI 2025 · 3 citations
- Engaging Communities Meaningfully in Defining Disability Representation for AI Image GenerationAnja Thieme, Rita Faia Marques, Martin Grayson, Sidhika Balachandar et al.CHI 2026 · 1 citation
- From Reflection to Repair: A Scoping Review of Dataset Documentation ToolsPedro Reynolds-Cuéllar, Marisol Wong-Villacres, Adriana Alvarado Garcia, Heila PrecelCHI 2026 · 1 citation
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Towards Understanding and Mitigating Social Biases in Language ModelsPaul Pu Liang, Chiyu Wu, Louis-Philippe Morency, Ruslan SalakhutdinovICML 2021 · 495 citations
- Bias Out-of-the-Box: An Empirical Analysis of Intersectional Occupational Biases in Popular Generative Language ModelsHannah Rose Kirk, Yennie Jun, Filippo Volpin, Haider Iqbal et al.NeurIPS 2021 · 243 citations
- Interpreting Pretrained Contextualized Representations via Reductions to Static EmbeddingsRishi Bommasani, Kelly Davis, Claire CardieACL 2020 · 137 citations
- Visual Analysis of Discrimination in Machine LearningQianwen Wang, Zhenhua Xu, Chen Zhu-Tian, Yong Wang et al.IEEE VIS 2020 · 56 citations
Related papers
- Debiasing Pretrained Text Encoders by Paying Attention to Paying AttentionYacine Gaci, Boualem Benatallah, Fabio Casati, Khalid BenabdeslemEMNLP 2022 · 12 citations
- Towards Debiasing Sentence RepresentationsPaul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim et al.ACL 2020 · 149 citations
- Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant LearningFan Zhou, Yuzhou Mao, Liu Yu, Yi Yang et al.ACL 2023 · 21 citations
- On Measuring and Mitigating Biased Inferences of Word EmbeddingsSunipa Dev, Tao Li, Jeff M. Phillips, Vivek SrikumarAAAI 2020 · 195 citations
- Job Unfair: An Investigation of Gender and Occupational Bias in Free-Form Text Completions by LLMsCamilla Casula, Sebastiano Vecellio Salto, Elisa Leonardelli, Sara TonelliEMNLP 2025
