Conversation for Non-verifiable Learning: Self-Evolving Large Language Models through Meta-Evaluation
Yuan Sui, Bryan Hooi
Abstract
Training large language models (LLMs) for nonverifiable tasks-such as creative writing, dialogue, and ethical reasoning-remains challenging due to the absence of ground-truth labels. While LLM-as-Judge approaches offer a scalable alternative to human feedback, they face a fundamental limitation: performance is constrained by the evaluator's own quality. If the judge cannot recognize good solutions, it cannot provide useful training signals, and evaluation biases (e.g., favoring verbosity over quality) remain unaddressed. This motivates meta-evaluation-the ability to evaluate and improve the evaluator itself. We introduce CoNL, a framework that unifies generation, evaluation, and meta-evaluation through multi-agent self-play. Our key insight: critique quality can be measured by whether it helps others improve their solutions. In CoNL, multiple agents sharing the same policy engage in structured conversations to propose, critique, and revise solutions. Critiques that enable other agents' solution improvements earn a diagnostic reward, creating explicit supervision for meta-evaluation and enabling joint optimization of generation and judging capabilities through self-play, without external judges or ground truth. Experiments on various benchmarks show that CoNL achieves consistent improvements over self-rewarding baselines while maintaining stable training. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5a7b13e-49aa-43ed-8822-a9be907f4c81Builds on12
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li et al.ICML 2024 · 569 citations
- Reinforcement Learning for Reasoning in Large Language Models with One Training ExampleYiping Wang, Qing Yang, Zhiyuan Zeng, Liliang Ren et al.NeurIPS 2025 · 314 citations
Related papers
- Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-JudgeTianhao Wu, Weizhe Yuan, Olga Golovneva, Jing Xu et al.EMNLP 2025 · 6 citations
- Teaching Language Models to Critique via Reinforcement LearningZhihui Xie, Jie Chen, Liyu Chen, Weichao Mao et al.ICML 2025
- J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement LearningChenxi Whitehouse, Tianlu Wang, Ping Yu, Xian Li et al.ICLR 2026 · 74 citations
- Teaching Models to Improve on TapeLiat Bezalel, Eyal Orgad, Amir GlobersonAAAI 2025
- Interactive Evaluation of Large Language Models for Multi-Requirement Software Engineering TasksDimitrios Rontogiannis, Maxime Peyrard, Nicolas Mario Baldwin, Martin Josifoski et al.AAAI 2026 · 1 citation
