Dataset Cartography: Mapping and Diagnosing Datasets with Training Dynamics
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, Yejin Choi
Abstract
Large datasets have become commonplace in NLP research. However, the increased emphasis on data quantity has made it challenging to assess the quality of data. We introduce Data Maps-a model-based tool to characterize and diagnose datasets. We leverage a largely ignored source of information: the behavior of the model on individual instances during training (training dynamics) for building data maps. This yields two intuitive measures for each example-the model's confidence in the true class, and the variability of this confidence across epochs-obtained in a single run of training. Experiments across four datasets show that these model-dependent measures reveal three distinct regions in the data map, each with pronounced characteristics. First, our data maps show the presence of ambiguous regions with respect to the model, which contribute the most towards out-of-distribution generalization. Second, the most populous regions in the data are easy to learn for the model, and play an important role in model optimization. Finally, data maps uncover a region with instances that the model finds hard to learn; these often correspond to labeling errors. Our results indicate that a shift in focus from quantity to quality of data could lead to robust models and improved out-ofdistribution generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 106d4606-e7f6-4f5e-ab89-4d755d7245d1Cited by top-tier papers139
- Identifying Mislabeled Data using the Area Under the Margin RankingGeoff Pleiss, Tianyi Zhang, Ethan R. Elenberg, Kilian Q. WeinbergerNeurIPS 2020 · 398 citations
- Understanding Dataset Difficulty with V-Usable InformationKawin Ethayarajh, Yejin Choi, Swabha SwayamdiptaICML 2022 · 337 citations
- SOTOPIA: Interactive Evaluation for Social Intelligence in Language AgentsXuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang et al.ICLR 2024 · 288 citations
- NumGLUE: A Suite of Fundamental yet Challenging Mathematical Reasoning TasksSwaroop Mishra, Arindam Mitra, Neeraj Varshney, Bhavdeep Singh Sachdeva et al.ACL 2022 · 138 citations
- Easy-to-Hard Generalization: Scalable Alignment Beyond Human SupervisionZhiqing Sun, Longhui Yu, Yikang Shen, Weiyang Liu et al.NeurIPS 2024 · 125 citations
Builds on5
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- Identifying Mislabeled Data using the Area Under the Margin RankingGeoff Pleiss, Tianyi Zhang, Ethan R. Elenberg, Kilian Q. WeinbergerNeurIPS 2020 · 398 citations
- Adversarial Filters of Dataset BiasesRonan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers et al.ICML 2020 · 242 citations
- Continual Deep Learning by Functional Regularisation of Memorable PastPingbo Pan, Siddharth Swaroop, Alexander Immer, Runa Eschenhagen et al.NeurIPS 2020 · 179 citations
- Simple and Effective Regularization Methods for Training on Noisily Labeled Data with Generalization GuaranteeWei Hu, Zhiyuan Li, Dingli YuICLR 2020 · 140 citations
Related papers
- Your Model is Overconfident, and Other Lies We Tell OurselvesTimothee Mickus, Aman Sinha, Raúl VázquezACL 2025
- When Annotators Disagree, Topology Explains: Mapper, a Topological Tool for Exploring Text Embedding Geometry and AmbiguityNisrine Rair, Alban Goupil, Valeriu Vrabie, Emmanuel ChochoyEMNLP 2025
- Instance-adaptive training with noise-robust losses against noisy labelsLifeng Jin, Linfeng Song, Kun Xu, Dong YuEMNLP 2021 · 9 citations
- Data-SUITE: Data-centric identification of in-distribution incongruous examplesNabeel Seedat, Jonathan Crabbé, Mihaela van der SchaarICML 2022 · 16 citations
- DMAP: A Distribution Map for TextTom Kempton, Julia Rozanova, Parameswaran Kamalaruban, Maeve Madigan et al.ICLR 2026 · 1 citation
