RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
Mian Wu, Gavin Zhang, Sewon Min, Sergey Levine, Aviral Kumar
Abstract
Open-ended generation tasks require outputs to satisfy diverse and often implicit task-specific evaluation rubrics. The sheer number of relevant rubrics leads to prohibitively high verification costs and incomplete assessments of a response, making reinforcement learning (RL) post-training with rubric-based rewards difficult to scale. This problem is exacerbated by the fact that often the best way to combine these rubrics into one single reward is also highly prompt-specific. We propose Reinforcement Learning with Adversarial Critic (RLAC), a post-training approach that addresses these challenges via dynamic rubric verification. Our approach employs a large language model (LLM) as a critic that dynamically identifies only the most likely failure modes (e.g., a factual error or unhandled edge case), which are then verified by an external validator to optimize both generator and critic jointly. By training both the generator and the critic, this game enhances the critic's error detection and the generator's output quality while reducing required verifications. Our experiments demonstrate that RLAC improves factual accuracy in text generation and correctness in code generation, while also outperforming exhaustive verification and reward model methods. We show that dynamic critics are more effective than fixed critics, showcasing the potential of RLAC for scaling RL post-training to free-form generation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cff8bbe7-2019-45f9-994c-3d5f2e4e2eacCited by top-tier papers4
- Reinforcement Learning with Evolving Rubrics for Deep ResearchRulin Shao, Akari Asai, Shannon Shen, Hamish Ivison et al.ICML 2026 · 78 citations
- Agentic Rubrics as Contextual Verifiers for SWE AgentsMohit Raghavendra, Anisha Gunjal, Bing Liu, Yunzhong HeACL 2026 · 10 citations
- Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents with Citation-Aware Rubric RewardsJiajie Zhang, Xin Lv, Ling Feng, Lei Hou et al.ACL 2026 · 8 citations
- Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement LearningWeiqin Wang, Yile Wang, Kehao Chen, Hui HuangACL 2026 · 5 citations
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
Related papers
- ATGen: Adversarial Reinforcement Learning for Test Case GenerationQingyao Li, Xinyi Dai, Weiwen Liu, Xiangyang Li et al.ICLR 2026 · 4 citations
- Online Rubrics Elicitation from Pairwise ComparisonsMohammadHossein Rezaei, Robert Vacareanu, Zihao Wang, Clinton Wang et al.ICML 2026 · 41 citations
- Teaching Language Models to Critique via Reinforcement LearningZhihui Xie, Jie Chen, Liyu Chen, Weichao Mao et al.ICML 2025
- Crossing the Reward Bridge: Expanding Reinforcement Learning with Verifiable Rewards Across Diverse DomainsYi Su, Dian Yu, Linfeng Song, Juntao Li et al.ACL 2026
- Escaping the Verifier: Learning to Reason via DemonstrationsLocke Cai, Max Ryabinin, Ivan ProvilkovICML 2026 · 6 citations
