Changing the World by Changing the Data
Anna Rogers
摘要
NLP community is currently investing a lot more research and resources into development of deep learning models than training data. While we have made a lot of progress, it is now clear that our models learn all kinds of spurious patterns, social biases, and annotation artifacts. Algorithmic solutions have so far had limited success. An alternative that is being actively discussed is more careful design of datasets so as to deliver specific signals. This position paper maps out the arguments for and against data curation, and argues that fundamentally the point is moot: curation already is and will be happening, and it is changing the world. The question is only how much thought we want to invest into that process.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
- Jury Learning: Integrating Dissenting Voices into Machine Learning ModelsMitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel 等CHI 2022 · 被引用 134 次
- NLPositionality: Characterizing Design Biases of Datasets and ModelsSebastin Santy, Jenny T. Liang, Ronan Le Bras, Katharina Reinecke 等ACL 2023 · 被引用 23 次
- BLIP: Facilitating the Exploration of Undesirable Consequences of Digital TechnologiesRock Yuren Pang, Sebastin Santy, René Just, Katharina ReineckeCHI 2024 · 被引用 19 次
- ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language ModelsAparna Elangovan, Ling Liu, Lei Xu, Sravan Babu Bodapati 等ACL 2024 · 被引用 19 次
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Climbing towards NLU: On Meaning, Form, and Understanding in the Age of DataEmily M. Bender, Alexander KollerACL 2020 · 被引用 914 次
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong 等CHI 2021 · 被引用 725 次
- Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image RepresentationsTianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang 等ICCV 2019 · 被引用 469 次
- Adversarial Watermarking Transformer: Towards Tracing Text Provenance with Data HidingSahar Abdelnabi, Mario FritzS&P 2021 · 被引用 210 次
相关 Paper
- Escaping Collapse: The Strength of Weak Data for Large Language Model TrainingKareem Amin, Sara Babakniya, Alex Bie, Weiwei Kong 等NeurIPS 2025 · 被引用 17 次
- Integrating Machine Learning Data with Symbolic Knowledge from Collaboration Practices of Curators to Improve Conversational SystemsClaudio Santos Pinhanez, Heloisa Candello, Paulo Rodrigo Cavalin, Mauro Carlos Pichiliani 等CHI 2021 · 被引用 7 次
- Explaining the Efficacy of Counterfactually Augmented DataDivyansh Kaushik, Amrith Setlur, Eduard H. Hovy, Zachary Chase LiptonICLR 2021 · 被引用 89 次
- Thinking beyond the anthropomorphic paradigm benefits LLM researchLujain Ibrahim, Myra ChengACL 2026 · 被引用 13 次
- Exploring the Efficacy of Automatically Generated Counterfactuals for Sentiment AnalysisLinyi Yang, Jiazheng Li, Padraig Cunningham, Yue Zhang 等ACL 2021
