Stepmothers are mean and academics are pretentious: What do pretrained language models learn about you?
Rochelle Choenni, Ekaterina Shutova, Robert van Rooij
摘要
In this paper, we investigate what types of stereotypical information are captured by pretrained language models. We present the first dataset comprising stereotypical attributes of a range of social groups and propose a method to elicit stereotypes encoded by pretrained language models in an unsupervised fashion. Moreover, we link the emergent stereotypes to their manifestation as basic emotions as a means to study their emotional effects in a more generalized manner. To demonstrate how our methods can be used to analyze emotion and stereotype shifts due to linguistic experience, we use fine-tuning on news sources as a case study. Our experiments expose how attitudes towards different social groups vary across models and how quickly emotions and stereotypes can shift at the fine-tuning stage.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Deciphering Stereotypes in Pre-Trained Language ModelsWeicheng Ma, Henry Scheible, Brian Wang, Goutham Veeramachaneni 等EMNLP 2023 · 被引用 7 次
- KidLM: Advancing Language Models for Children - Early Insights and Future DirectionsMir Tafseer Nayeem, Davood RafieiEMNLP 2024 · 被引用 7 次
- A Comprehensive Framework to Operationalize Social Stereotypes for Responsible AI EvaluationsAida Mostafazadeh Davani, Sunipa Dev, Héctor Pérez-Urbina, Vinodkumar PrabhakaranEMNLP 2025 · 被引用 6 次
- The Echoes of Multilinguality: Tracing Cultural Value Shifts during Language Model Fine-tuningRochelle Choenni, Anne Lauscher, Ekaterina ShutovaACL 2024 · 被引用 2 次
它引用的顶会 Paper5
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 被引用 68 次
- CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language ModelsNikita Nangia, Clara Vania, Rasika Bhalerao, Samuel R. BowmanEMNLP 2020 · 被引用 19 次
- Unsupervised Discovery of Implicit Gender BiasAnjalie Field, Yulia TsvetkovEMNLP 2020 · 被引用 5 次
- StereoSet: Measuring stereotypical bias in pretrained language modelsMoin Nadeem, Anna Bethke, Siva ReddyACL 2021
相关 Paper
- Upstream Mitigation Is Not All You Need: Testing the Bias Transfer Hypothesis in Pre-Trained Language ModelsRyan Steed, Swetasudha Panda, Ari Kobren, Michael L. WickACL 2022 · 被引用 52 次
- Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant LearningFan Zhou, Yuzhou Mao, Liu Yu, Yi Yang 等ACL 2023 · 被引用 21 次
- Angry Men, Sad Women: Large Language Models Reflect Gendered Stereotypes in Emotion AttributionFlor Miriam Plaza del Arco, Amanda Cercas Curry, Alba Cercas Curry, Gavin Abercrombie 等ACL 2024 · 被引用 8 次
- Does the Emotional Understanding of LVLMs Vary Under High-Stress Environments and Across Different Demographic Attributes?Jaewook Lee, Yeajin Jang, Oh-Woog Kwon, Harksoo KimACL 2025
- From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP ModelsShangbin Feng, Chan Young Park, Yuhan Liu, Yulia TsvetkovACL 2023 · 被引用 117 次
