Lune

ACL2023Top-tier venue

Targeted Data Generation: Finding and Fixing Model Weaknesses

Zexue He, Marco Túlio Ribeiro, Fereshte Khani

2023Year
6Citations
8Top-tier citations

Abstract

Even when aggregate accuracy is high, stateof-the-art NLP models often fail systematically on specific subgroups of data, resulting in unfair outcomes and eroding user trust. Additional data collection may not help in addressing these weaknesses, as such challenging subgroups may be unknown to users, and underrepresented in the existing and new data. We propose Targeted Data Generation (TDG), a framework that automatically identifies challenging subgroups, and generates new data for those subgroups using large language models (LLMs) with a human in the loop. TDG estimates the expected benefit and potential harm of data augmentation for each subgroup, and selects the ones most likely to improve withingroup performance without hurting overall performance. In our experiments, TDG 1 significantly improves the accuracy on challenging subgroups for state-of-the-art sentiment analysis and natural language inference models, while also improving overall test accuracy. * Work done during the internship at Microsoft. 1 Codes and collected data will be released in https:// github.com/ZexueHe/TDG .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers8

Ask how each one uses it

Builds on7

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines