Quantifying Uncertainty in Error Consistency: Towards Reliable Behavioral Comparison of Classifiers
Thomas Klein, Sascha Meyen, Wieland Brendel, Felix A. Wichmann, Kristof Meding
Abstract
Benchmarking models is a key factor for the rapid progress in machine learning (ML) research. Thus, further progress depends on improving benchmarking metrics. A standard metric to measure the behavioral alignment between ML models and human observers is error consistency (EC). EC allows for more fine-grained comparisons of behavior than other metrics such as accuracy, and has been used in the influential Brain-Score benchmark to rank different DNNs by their behavioral consistency with humans. Previously, EC values have been reported without confidence intervals. However, empirically measured EC values are typically noisy -- thus, without confidence intervals, valid benchmarking conclusions are problematic. Here we improve on standard EC in two ways: First, we show how to obtain confidence intervals for EC using a bootstrapping technique, allowing us to derive significance tests for EC. Second, we propose a new computational model relating the EC between two classifiers to the implicit probability that one of them copies responses from the other. This view of EC allows us to give practical guidance to scientists regarding the number of trials required for sufficiently powerful, conclusive experiments. Finally, we use our methodology to revisit popular NeuroAI-results. We find that while the general trend of behavioral differences between humans and machines holds up to scrutiny, many reported differences between deep vision models are statistically insignificant. Our methodology enables researchers to design adequately powered experiments that can reliably detect behavioral differences between models, providing a foundation for more rigorous benchmarking of behavioral alignment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf6e064d-0349-4884-950a-cda6825ad77bCited by top-tier papers1
Ask how each one uses itBuilds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Partial success in closing the gap between human and machine visionRobert Geirhos, Kantharaju Narayanappa, Benjamin Mitzkus, Tizian Thieringer et al.NeurIPS 2021 · 304 citations
- Beyond accuracy: quantifying trial-by-trial behaviour of CNNs and humans by measuring error consistencyRobert Geirhos, Kristof Meding, Felix A. WichmannNeurIPS 2020 · 154 citations
- Self-supervised video pretraining yields robust and more human-aligned visual representationsNikhil Parthasarathy, S. M. Ali Eslami, João Carreira, Olivier J. HénaffNeurIPS 2023 · 27 citations
- When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model LeaderboardsNorah A. Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed et al.ACL 2024 · 13 citations
Related papers
- Only Brains Align with Brains: Cross-Region Alignment Patterns Expose Limits of Normative ModelsLarissa Höfling, Matthias Tangemann, Lotta Piefke, Susanne Keller et al.ICLR 2026 · 5 citations
- Human alignment of neural network representationsLukas Muttenthaler, Jonas Dippel, Lorenz Linhardt, Robert A. Vandermeulen et al.ICLR 2023 · 15 citations
- Model-Behavior Alignment under Flexible Evaluation: When the Best-Fitting Model Isn't the Right OneItamar Avitan, Tal GolanNeurIPS 2025 · 5 citations
- One Hundred Neural Networks and Brains Watching Videos: Lessons from AlignmentChristina Sartzetaki, Gemma Roig, Cees G. M. Snoek, Iris I. A. GroenICLR 2025
- Generalized Shape Metrics on Neural RepresentationsAlex H. Williams, Erin Kunz, Simon Kornblith, Scott W. LindermanNeurIPS 2021 · 182 citations
