Your Model is Overconfident, and Other Lies We Tell Ourselves
Timothee Mickus, Aman Sinha, Raúl Vázquez
Abstract
The difficulty intrinsic to a given example, rooted in its inherent ambiguity, is a key yet often overlooked factor in evaluating neural NLP models. We investigate the interplay and divergence among various metrics for assessing intrinsic difficulty, including annotator dissensus, training dynamics, and model confidence. Through a comprehensive analysis using 29 models on three datasets, we reveal that while correlations exist among these metrics, their relationships are neither linear nor monotonic. By disentangling these dimensions of uncertainty, we aim to refine our understanding of data complexity and its implications for evaluating and improving NLP models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c48fb5e6-1746-4d18-a18b-d57be0a8bd84Builds on9
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 439 citations
- Human Uncertainty Makes Classification More RobustJoshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths, Olga RussakovskyICCV 2019 · 362 citations
- Deep Learning Through the Lens of Example DifficultyRobert J. N. Baldock, Hartmut Maennel, Behnam NeyshaburNeurIPS 2021 · 204 citations
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpaceMor Geva, Avi Caciularu, Kevin Ro Wang, Yoav GoldbergEMNLP 2022 · 92 citations
Related papers
- Dataset Cartography: Mapping and Diagnosing Datasets with Training DynamicsSwabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang et al.EMNLP 2020 · 12 citations
- Rethinking Code Complexity Through the Lens of Large Language ModelsChen Xie, Xiaodong Gu, Yuling Shi, Beijun ShenICML 2026
- Intrinsic Evaluation of Summarization DatasetsRishi Bommasani, Claire CardieEMNLP 2020 · 52 citations
- Modeling Disclosive Transparency in NLP Application DescriptionsMichael Saxon, Sharon Levy, Xinyi Wang, Alon Albalak et al.EMNLP 2021 · 3 citations
- We're Afraid Language Models Aren't Modeling AmbiguityAlisa Liu, Zhaofeng Wu, Julian Michael, Alane Suhr et al.EMNLP 2023 · 35 citations
