Lune

ACL2025Top-tier venue

Your Model is Overconfident, and Other Lies We Tell Ourselves

Timothee Mickus, Aman Sinha, Raúl Vázquez

2025Year

Abstract

The difficulty intrinsic to a given example, rooted in its inherent ambiguity, is a key yet often overlooked factor in evaluating neural NLP models. We investigate the interplay and divergence among various metrics for assessing intrinsic difficulty, including annotator dissensus, training dynamics, and model confidence. Through a comprehensive analysis using 29 models on three datasets, we reveal that while correlations exist among these metrics, their relationships are neither linear nor monotonic. By disentangling these dimensions of uncertainty, we aim to refine our understanding of data complexity and its implications for evaluating and improving NLP models.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext c48fb5e6-1746-4d18-a18b-d57be0a8bd84

Builds on9

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines