Targeted Data Generation: Finding and Fixing Model Weaknesses
Zexue He, Marco Túlio Ribeiro, Fereshte Khani
摘要
Even when aggregate accuracy is high, stateof-the-art NLP models often fail systematically on specific subgroups of data, resulting in unfair outcomes and eroding user trust. Additional data collection may not help in addressing these weaknesses, as such challenging subgroups may be unknown to users, and underrepresented in the existing and new data. We propose Targeted Data Generation (TDG), a framework that automatically identifies challenging subgroups, and generates new data for those subgroups using large language models (LLMs) with a human in the loop. TDG estimates the expected benefit and potential harm of data augmentation for each subgroup, and selects the ones most likely to improve withingroup performance without hurting overall performance. In our experiments, TDG 1 significantly improves the accuracy on challenging subgroups for state-of-the-art sentiment analysis and natural language inference models, while also improving overall test accuracy. * Work done during the internship at Microsoft. 1 Codes and collected data will be released in https:// github.com/ZexueHe/TDG .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- MedEval: A Multi-Level, Multi-Task, and Multi-Domain Medical Benchmark for Language Model EvaluationZexue He, Yu Wang, An Yan, Yao Liu 等EMNLP 2023 · 被引用 7 次
- LLM-enhanced Self-training for Cross-domain Constituency ParsingJianling Li, Meishan Zhang, Peiming Guo, Min Zhang 等EMNLP 2023 · 被引用 3 次
- What Is Wrong with My Model? Identifying Systematic Problems with Semantic Data SlicingChenyang Yang, Yining Hong, Grace A. Lewis, Tongshuang Wu 等ASE 2024 · 被引用 2 次
- Targeted Distillation for Sentiment AnalysisYice Zhang, Guangyu Xie, Jingjie Lin, Jianzhu Bao 等EMNLP 2025 · 被引用 2 次
- Composition-Grounded Data Synthesis for Visual ReasoningXinyi Gu, Jiayuan Mao, Zhang-Wei Hong, Zhuoran Yu 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Do Not Have Enough Data? Deep Learning to the Rescue!Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor 等AAAI 2020 · 被引用 398 次
- No Subclass Left Behind: Fine-Grained Robustness in Coarse-Grained Classification ProblemsNimit Sharad Sohoni, Jared Dunnmon, Geoffrey Angus, Albert Gu 等NeurIPS 2020 · 被引用 316 次
- Robustness to Spurious Correlations via Human AnnotationsMegha Srivastava, Tatsunori B. Hashimoto, Percy LiangICML 2020 · 被引用 103 次
- Adaptive Testing and Debugging of NLP ModelsMarco Túlio Ribeiro, Scott M. LundbergACL 2022 · 被引用 99 次
相关 Paper
- BTC-SAM: Leveraging LLMs for Generation of Bias Test Cases for Sentiment Analysis ModelsZsolt T. Kardkovács, Lynda Djennane, Anna Field, Boualem Benatallah 等EMNLP 2025
- Exploring the Efficacy of Automatically Generated Counterfactuals for Sentiment AnalysisLinyi Yang, Jiazheng Li, Padraig Cunningham, Yue Zhang 等ACL 2021
- Generating Data for Symbolic Language with Large Language ModelsJiacheng Ye, Chengzu Li, Lingpeng Kong, Tao YuEMNLP 2023 · 被引用 8 次
- ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech DetectionThomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap 等ACL 2022
- Can You Rely on Your Model Evaluation? Improving Model Evaluation with Synthetic Test DataBoris van Breugel, Nabeel Seedat, Fergus Imrie, Mihaela van der SchaarNeurIPS 2023 · 被引用 51 次
