InsideOut: Measuring and Mitigating Insider-Outsider Bias in Interview Script Generation
Yixin Wan, Xingrun Chen, Kai-Wei Chang
Abstract
Advancements in Large language models (LLMs) have enabled a variety of downstream applications like story and interview script generation. However, recent research raised concerns about culture-related fairness issues in LLM-generated content. In this work, we identify and systematically investigate LLMs' insider-outsider bias, a phenomenon where models position themselves as "insiders" of mainstream cultures during generation while externalizing less dominant cultures. We propose the INSIDEOUT benchmark with 4,000 generation prompts and three evaluation metrics to quantify this bias through a culturally situated interview script generation task, in which an LLM is positioned as a reporter interviewing local people across 10 diverse cultures. Empirical evaluation on 5 state-of-the-art LLMs reveals that while models adopt insider tones in over 88% US-contexted scripts on average, they disproportionately default to "outsider" stances for non-Western cultures. To mitigate these biases, we propose 2 inference-time methods: a baseline prompt-based Fairness Intervention Pillars (FIP) method, and a structured Mitigation via Fairness Agents (MFA) framework consisting of a Single-Agent (MFA-SA), a Hierarchical-Agent (MFA-HA), and an autonomous Agentic Planning (MFA-Plan) pipeline. Empirical results demonstrate that agent-based MFA methods achieve outstanding and robust performance in mitigating the insider-outsider bias: For instance, on the Cultural Alignment Gap (CAG) metric, MFA-SA reduces bias in Llama model by 89.70 % and MFA-HA mitigates bias in Qwen by 82.54%. These findings showcase the effectiveness of agent-based methods as a promising direction for mitigating biases in generative LLMs. * Equal contribution. US Culture Non-US Cultures LLM LLM You sound like an "insider" in this culture Cultural Perspective Deviation (CPD): 20.00 Cultural Alignment Gap (CAG): 55.00 Ho cu ap en Ho tra sh Ch Ca ho int Ch op Ho yo en inf va Fin yo en int th Can you tell us more about your current venture and what inspired you to start it? In your experience, what are some of the unique challenges that small businesses face in the United States today? How do you think technology can help address these issues? How do you see the future of entrepreneurship in the United States evolving over the next few years? Lastly, what advice would you give to aspiring entrepreneurs who are looking to start their own businesses? United States You sound like an "outsider" in this culture How do you think the current emphasis on diversity and inclusion is influencing American culture? Can you share some aspects of your culture that you're particularly proud of? "You are a journalist conducting on-the-ground interviews about ___ culture..."
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human PreferencesShreya Shankar, J. D. Zamfirescu-Pereira, Bjoern Hartmann, Aditya G. Parameswaran et al.UIST 2024 · 143 citations
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
- Investigating Cultural Alignment of Large Language ModelsBadr AlKhamissi, Muhammad N. ElNokrashy, Mai Alkhamissi, Mona T. DiabACL 2024 · 27 citations
- White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMsYixin Wan, Kai-Wei ChangACL 2025 · 9 citations
Related papers
- A Game-Theoretica Negotiation Framework for Cross-Cultural ConsensusGuoxi Zhang, Jiawei Chen, Tianzhuo Yang, Jiaming Ji et al.ACL 2026
- From Single to Societal: Analyzing Persona-Induced Bias in Multi-Agent InteractionsJiayi Li, Xiao Liu, Yansong FengAAAI 2026 · 3 citations
- Unbiased Evaluation of Large Language Models from a Causal PerspectiveMeilin Chen, Jian Tian, Liang Ma, Di Xie et al.ICML 2025
- FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMsZhiting Fan, Ruizhe Chen, Tianxiang Hu, Zuozhu LiuICLR 2025
- F²Bench: An Open-ended Fairness Evaluation Benchmark for LLMs with Factuality ConsiderationsTian Lan, Jiang Li, Yemin Wang, Xu Liu et al.EMNLP 2025 · 3 citations
