Correlated Errors in Large Language Models
Elliot Myunghoon Kim, Avi Garg, Kenny Peng, Nikhil Garg
Abstract
Diversity in training data, architecture, and providers is assumed to mitigate homogeneity in LLMs. However, we lack empirical evidence on whether different LLMs differ meaningfully. We conduct a large-scale empirical evaluation on over 350 LLMs overall, using two popular leaderboards and a resume-screening task. We find substantial correlation in model errors-on one leaderboard dataset, models agree 60% of the time when both models err. We identify factors driving model correlation, including shared architectures and providers. Crucially, however, larger and more accurate models have highly correlated errors, even with distinct architectures and providers. Finally, we show the effects of correlation in two downstream tasks: LLM-as-judge evaluation and hiring-the latter reflecting theoretical predictions regarding algorithmic monoculture.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3700b6b-241a-46a7-80a8-cfdbfbe8b8a6Cited by top-tier papers12
- Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment DatasetLily H Zhang, Smitha Milli, Karen Long Jusko, Jonathan Smith et al.ICLR 2026 · 41 citations
- The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?Alexander Hägele, Aryo Pradipta Gema, Henry Sleight, Ethan Perez et al.ICLR 2026 · 10 citations
- PolicyPad: Collaborative Prototyping of LLM PoliciesK. J. Kevin Feng, Tzu-Sheng Kuo, Quan Ze Jim Chen, Inyoung Cheong et al.CHI 2026 · 2 citations
- Cutting LLM Evaluation Costs with SySRs: A Bandit Algorithm That Provably Exploits Model SimilarityZifan Lyu, Chahine Nejma, Tobias Wegel, Fanny Yang et al.ICML 2026
- Truthfulness Does Not Scale Like Reasoning: Why Polling Fails as a Proxy VerifierYegor Denisov-Blanch, Joshua Kazdan, Jessica Chudnovsky, Rylan Schaeffer et al.ICML 2026
Builds on10
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 865 citations
- Picking on the Same Person: Does Algorithmic Monoculture lead to Outcome Homogenization?Rishi Bommasani, Kathleen A. Creel, Ananya Kumar, Dan Jurafsky et al.NeurIPS 2022 · 179 citations
- Beyond accuracy: quantifying trial-by-trial behaviour of CNNs and humans by measuring error consistencyRobert Geirhos, Kristof Meding, Felix A. WichmannNeurIPS 2020 · 154 citations
- Harnessing the Universal Geometry of EmbeddingsRishi D. Jha, Collin Zhang, Vitaly Shmatikov, John X. MorrisNeurIPS 2025 · 69 citations
Related papers
- Monoculture or Multiplicity: Which Is It?Mila Gorecki, Moritz HardtNeurIPS 2025 · 6 citations
- CARE: Confounder-Aware Aggregation for Reliable LLM EvaluationJitian Zhao, Changho Shin, Tzu-Heng Huang, Satya Sai Srinath Namburi GNVV et al.ICML 2026 · 7 citations
- Generative Monoculture in Large Language ModelsFan Wu, Emily Black, Varun ChandrasekaranICLR 2025
- When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model LeaderboardsNorah A. Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed et al.ACL 2024 · 13 citations
- AI Cartography: Mapping the Latent Landscape of AI Benchmark EcosystemsMichael Hardy, Anka Reuel, Lijin Zhang, Jodi Casabianca et al.ICML 2026
