SCoUT: A Framework for Structured Stereotype Analysis in Language Models
Jinxuan Wu, Bin Li, Xiangyang Xue
摘要
Existing stereotype auditing methods for Large Language Models (LLMs) typically rely on isolated rating schemes or task-specific probes, lacking theoretical grounding and failing to reveal internal organization beyond surface-level output patterns. In this paper, we introduce SCoUT (Stereotype Content-oriented Utility structure via Thurstonian modeling), a closed-loop framework that structurally models, explicitly probes, and functionally steers stereotype dimensions (warmth and competence) in LLMs. SCoUT first reconstructs a global stereotype utility structure aligned with Stereotype Content Model theory via Thurstonian comparative judgments. Across multiple open-source LLMs, this modeling achieves high pairwise-preference prediction accuracy (≥ 0.90 on larger-scale models) and exhibits strong cross-model consistency. Probing internal attention mechanisms localizes this structure to specific heads (Spearman’s ρ up to 0.83 for warmth and 0.90 for competence) and surfaces a salient asymmetry between warmth and competence. Further, targeted inference-time activation modifications on these dimension-sensitive heads consistently steer model outputs along the intended axes. By bridging behavioral measurement with internal representation and controllable steering, SCoUT offers an end-to-end framework that uncovers and interprets the latent structure of stereotypes, advancing stereotype auditing from surface detection to structural analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language ModelsNikita Nangia, Clara Vania, Rasika Bhalerao, Samuel R. BowmanEMNLP 2020 · 被引用 19 次
- StereoMap: Quantifying the Awareness of Human-like Stereotypes in Large Language ModelsSullam Jeoung, Yubin Ge, Jana DiesnerEMNLP 2023 · 被引用 3 次
- Linear Representations of Political Perspective Emerge in Large Language ModelsJunsol Kim, James Evans, Aaron ScheinICLR 2025
- StereoSet: Measuring stereotypical bias in pretrained language modelsMoin Nadeem, Anna Bethke, Siva ReddyACL 2021
相关 Paper
- Understanding and Countering Stereotypes: A Computational Approach to the Stereotype Content ModelKathleen C. Fraser, Isar Nejadgholi, Svetlana KiritchenkoACL 2021
- Social-Group-Agnostic Bias Mitigation via the Stereotype Content ModelAli Omrani, Alireza Salkhordeh Ziabari, Charles Yu, Preni Golazizian 等ACL 2023 · 被引用 13 次
- Annotating Dimensions of Social Perception in Text: A Sentence-Level Dataset of Warmth and CompetenceMutaz Ayesh, Saif M. Mohammad, Nedjma OusidhoumACL 2026
- Uncovering Competency Gaps in Large Language Models and Their BenchmarksMaty Bohacek, Nino Scherrer, Nicholas Dufour, Thomas Leung 等ICML 2026
- The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM ObjectivesMatthieu Bou, Nyal Patel, Arjun Jagota, Satyapriya Krishna 等ICLR 2026 · 被引用 1 次
