Bad Seeds: Evaluating Lexical Methods for Bias Measurement
Maria Antoniak, David Mimno
摘要
Scope of Reproducibility In this work we verify the results of Bad Seeds: Evaluating Lexical Methods for Bias Measurement (Antoniak and Mimno, 2021). We replicate the experiments conducted and verify the main claims made in the original paper: (1) Bias measurements depend on seeds and models. (2) Shuffled seed pairs can result in a significant different bias subspace compared to ordered seed pairs. (3) Set similarity is negatively correlated with the explained variance of the first PCA component in the seed pairings subspace. Methodology We used skip-gram with negative sampling to train word2vec models with the same hyperparameters and data. We implemented code for the experiments using the resulting word embeddings and the seed sets provided by the authors. Results Overall, only one claim was reproduced. We reproduced the claim that bias measurements is dependent on the choice of seed set. We were not able to adequately reproduce the claims that shuffled pairs of seed sets generally result in less clearly defined correlation and that for pair of seed sets set similarity is negatively correlated with the explained variance of the first principal component. What was easy The paper is easy to follow. The data was publicly available. Authors replied frequently providing details about the parameters and preprocessing steps. Also authors were open for the discussion regarding seeds on the Github repository of the project. What was difficult In certain cases, the gathered seed sets json file contained errors. Specifically: 'daughters' was misspelled as 'daughers' (has now been updated). 'ma', 'am' was used as two words instead of one word "ma'am". The seed words for figure 4 as stated in the appendix are not the same as the one in the image itself. Table 2 from the original paper was also difficult to reproduce as the preprocessing according to the authors' description gave close but not equal results for the NYT dataset and significantly different results for two other datasets. Reproduced numbers are presented in table 3.3. Because of the time constraints, 20 bootstrapped launches were not conducted. Communication with original authors Contact was made with the original authors on multiple occasions to ask for clarification questions regarding the implementation of the experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Bias Out-of-the-Box: An Empirical Analysis of Intersectional Occupational Biases in Popular Generative Language ModelsHannah Rose Kirk, Yennie Jun, Filippo Volpin, Haider Iqbal 等NeurIPS 2021 · 被引用 243 次
- Picking on the Same Person: Does Algorithmic Monoculture lead to Outcome Homogenization?Rishi Bommasani, Kathleen A. Creel, Ananya Kumar, Dan Jurafsky 等NeurIPS 2022 · 被引用 179 次
- Faithful Explanations of Black-box NLP Models Using LLM-generated CounterfactualsYair Ori Gat, Nitay Calderon, Amir Feder, Alexander Chapanin 等ICLR 2024 · 被引用 55 次
- Rethinking "Risk" in Algorithmic Systems Through A Computational Narrative Analysis of Casenotes in Child-WelfareDevansh Saxena, Erina Seh-Young Moon, Aryan Chaurasia, Yixin Guan 等CHI 2023 · 被引用 35 次
- Under the Morphosyntactic Lens: A Multifaceted Evaluation of Gender Bias in Speech TranslationBeatrice Savoldi, Marco Gaido, Luisa Bentivogli, Matteo Negri 等ACL 2022 · 被引用 30 次
相关 Paper
- Assessing the Reliability of Word Embedding Gender Bias MeasuresYupei Du, Qixiang Fang, Dong NguyenEMNLP 2021 · 被引用 13 次
- Deep Fair Clustering for Visual LearningPeizhao Li, Han Zhao, Hongfu LiuCVPR 2020
- Intrinsic Bias Metrics Do Not Correlate with Application BiasSeraphina Goldfarb-Tarrant, Rebecca Marchant, Ricardo Muñoz Sánchez, Mugdha Pandya 等ACL 2021
- Reproducibility Issues for BERT-based Evaluation MetricsYanran Chen, Jonas Belouadi, Steffen EgerEMNLP 2022 · 被引用 11 次
- Towards Robustifying NLI Models Against Lexical Dataset BiasesXiang Zhou, Mohit BansalACL 2020 · 被引用 36 次
