SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
Dongwei Jiang, Jingyu Zhang, Orion Weller, Nathaniel Weir, Benjamin Van Durme, Daniel Khashabi
Abstract
Can LLMs consistently improve their previous outputs for better results? For this to be true, LLMs would need to be better at discriminating among previously-generated alternatives, than generating initial responses. We explore the validity of this hypothesis in practice. We first formulate a unified framework that allows us to compare the generative and discriminative capability of any model on any task. In our resulting experimental analysis of several open-source and industrial LLMs, we observe that model’s are not reliably better at discriminating among previously-generated alternatives than generating initial responses. This finding challenges the notion that LLMs may be able to enhance their performance only through their own judgment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dff58889-853d-447a-ab49-2c007cf6687bCited by top-tier papers4
- FEEDBACK FRICTION: LLMs Struggle to Fully Incorporate External FeedbackDongwei Jiang, Bowei Zhang, Andrew Wang, Nicholas Andrews et al.NeurIPS 2025 · 11 citations
- ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific DiscoveryZiru Chen, Shijie Chen, Yuting Ning, Qianheng Zhang et al.ICLR 2025 · 6 citations
- A Rolling Stone Gathers No Moss: Adaptive Policy Optimization for Stable Self-Evaluation in Large Multimodal ModelsWenkai Wang, Hongcan Guo, Zheqi Lv, Shengyu ZhangAAAI 2026
- Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained OptimizationTuan Nguyen, Long Tran-ThanhICML 2026
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson et al.ICML 2023 · 908 citations
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
Related papers
- IDGen: Item Discrimination Induced Prompt Generation for LLM EvaluationFan Lin, Shuyi Xie, Yong Dai, Wenlin Yao et al.NeurIPS 2024 · 7 citations
- The Benefits of Bad Advice: Autocontrastive Decoding across Model LayersAriel Gera, Roni Friedman, Ofir Arviv, R. Chulaka Gunasekara et al.ACL 2023 · 5 citations
- LLMmap: Fingerprinting for Large Language ModelsDario Pasquini, Evgenios M. Kornaropoulos, Giuseppe AtenieseUSENIX Security 2025
- Thinking LLMs: General Instruction Following with Thought GenerationTianhao Wu, Janice Lan, Weizhe Yuan, Jiantao Jiao et al.ICML 2025 · 3 citations
- Small Models Exhibit Limited Answer Consistency in Repetition Trials of the Multiple-Choice MMLU-Redux and MedQA BenchmarksClaudio S. Pinhanez, Paulo R. Cavalin, Cassia Sampaio Sanctos, Marcelo Carpinette GraveAAAI 2026
