Benchmarking and Improving Generator-Validator Consistency of Language Models
Xiang Lisa Li, Vaishnavi Shrivastava, Siyan Li, Tatsunori Hashimoto, Percy Liang
Abstract
As of September 2023, ChatGPT correctly answers "what is 7+8" with 15, but when asked "7+8=15, True or False" it responds with "False". This inconsistency between generating and validating an answer is prevalent in language models (LMs) and erodes trust. In this paper, we propose a framework for measuring the consistency between generation and validation (which we call generator-validator consistency, or GV-consistency), finding that even GPT-4, a state-of-the-art LM, is GV-consistent only 76% of the time. To improve the consistency of LMs, we propose to finetune on the filtered generator and validator responses that are GV-consistent, and call this approach consistency fine-tuning. We find that this approach improves GV-consistency of Alpaca-30B from 60% to 93%, and the improvement extrapolates to unseen tasks and domains (e.g., GV-consistency for positive style transfers extrapolates to unseen styles like humor). In addition to improving consistency, consistency fine-tuning improves both generator quality and validator accuracy without using any labeled data. Evaluated across 6 tasks, including math questions, knowledge-intensive QA, and instruction following, our method improves the generator quality by 16% and the validator accuracy by 6.3% across all tasks. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc8a933b-367c-4c31-ac63-8c678de30568Cited by top-tier papers15
- Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis RefinementLinlu Qiu, Liwei Jiang, Ximing Lu, Melanie Sclar et al.ICLR 2024 · 114 citations
- Unveiling the Tapestry of Consistency in Large Vision-Language ModelsYuan Zhang, Fei Xiao, Tao Huang, Chun-Kai Fan et al.NeurIPS 2024 · 27 citations
- Calibrating Verbalized Confidence with Self-Generated DistractorsVictor Wang, Elias Stengel-EskinICLR 2026 · 15 citations
- Precise Information Control in Long-Form Text GenerationJacqueline He, Howard Yen, Margaret Li, Shuyue Stella Li et al.NeurIPS 2025 · 8 citations
- On Evaluating LLM Alignment by Evaluating LLMs as JudgesYixin Liu, Pengfei Liu, Arman CohanNeurIPS 2025 · 7 citations
Builds on9
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai et al.EMNLP 2022 · 239 citations
- Maieutic Prompting: Logically Consistent Reasoning with Recursive ExplanationsJaehun Jung, Lianhui Qin, Sean Welleck, Faeze Brahman et al.EMNLP 2022 · 72 citations
Related papers
- Can Large Language Models Always Solve Easy Problems if They Can Solve Harder Ones?Zhe Yang, Yichang Zhang, Tianyu Liu, Jian Yang et al.EMNLP 2024 · 2 citations
- ConTested: Consistency-Aided Tested Code Generation with LLMJinhao Dong, Jun Sun, Wenjie Zhang, Jin Song Dong et al.ISSTA 2025 · 6 citations
- Consistency Analysis of ChatGPTMyeongjun Jang, Thomas LukasiewiczEMNLP 2023 · 55 citations
- The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to MistakesMarina Mancoridis, Zoe HitzigICML 2026
- Enhancing Self-Consistency and Performance of Pre-Trained Language Models through Natural Language InferenceEric Mitchell, Joseph J. Noh, Siyan Li, William S. Armstrong et al.EMNLP 2022 · 18 citations
