SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
Dongwei Jiang, Jingyu Zhang, Orion Weller, Nathaniel Weir, Benjamin Van Durme, Daniel Khashabi
摘要
Can LLMs consistently improve their previous outputs for better results? For this to be true, LLMs would need to be better at discriminating among previously-generated alternatives, than generating initial responses. We explore the validity of this hypothesis in practice. We first formulate a unified framework that allows us to compare the generative and discriminative capability of any model on any task. In our resulting experimental analysis of several open-source and industrial LLMs, we observe that model’s are not reliably better at discriminating among previously-generated alternatives than generating initial responses. This finding challenges the notion that LLMs may be able to enhance their performance only through their own judgment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- FEEDBACK FRICTION: LLMs Struggle to Fully Incorporate External FeedbackDongwei Jiang, Bowei Zhang, Andrew Wang, Nicholas Andrews 等NeurIPS 2025 · 被引用 11 次
- ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific DiscoveryZiru Chen, Shijie Chen, Yuting Ning, Qianheng Zhang 等ICLR 2025 · 被引用 6 次
- A Rolling Stone Gathers No Moss: Adaptive Policy Optimization for Stable Self-Evaluation in Large Multimodal ModelsWenkai Wang, Hongcan Guo, Zheqi Lv, Shengyu ZhangAAAI 2026
- Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained OptimizationTuan Nguyen, Long Tran-ThanhICML 2026
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson 等ICML 2023 · 被引用 908 次
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng 等ICLR 2024 · 被引用 858 次
相关 Paper
- IDGen: Item Discrimination Induced Prompt Generation for LLM EvaluationFan Lin, Shuyi Xie, Yong Dai, Wenlin Yao 等NeurIPS 2024 · 被引用 7 次
- The Benefits of Bad Advice: Autocontrastive Decoding across Model LayersAriel Gera, Roni Friedman, Ofir Arviv, R. Chulaka Gunasekara 等ACL 2023 · 被引用 5 次
- LLMmap: Fingerprinting for Large Language ModelsDario Pasquini, Evgenios M. Kornaropoulos, Giuseppe AtenieseUSENIX Security 2025
- Thinking LLMs: General Instruction Following with Thought GenerationTianhao Wu, Janice Lan, Weizhe Yuan, Jiantao Jiao 等ICML 2025 · 被引用 3 次
- Small Models Exhibit Limited Answer Consistency in Repetition Trials of the Multiple-Choice MMLU-Redux and MedQA BenchmarksClaudio S. Pinhanez, Paulo R. Cavalin, Cassia Sampaio Sanctos, Marcelo Carpinette GraveAAAI 2026
