CoAct: Co-Active LLM Preference Learning with Human-AI Synergy
Ruiyao Xu, Mihir Parmar, Tiankai Yang, Zhengyu Hu, Yue Zhao, Kaize Ding
Abstract
Learning from preference-based feedback has become an effective approach for aligning LLMs across diverse tasks. However, highquality human-annotated preference data remains expensive and scarce. Existing methods address this challenge through either selfrewarding, which scales by using purely AIgenerated labels but risks unreliability, or active learning, which ensures quality through oracle annotation but cannot fully leverage unlabeled data. In this paper, we present COACT, a novel framework that synergistically combines selfrewarding and active learning through strategic human-AI collaboration. COACT leverages self-consistency to identify both reliable selflabeled data and samples that are requiring oracle verification. Additionally, oracle feedback guides the model to generate new instructions within its solvable capability. Evaluated on three reasoning benchmarks across two model families, COACT achieves average improvements of +13.25% on GSM8K, +8.19% on MATH, and +13.16% on WebInstruct, consistently outperforming all baselines. 1 1 Our code is available at https://github.com/rux001/ CoAct .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a5506f0-7b29-4047-95f4-d6c790e63f0eBuilds on26
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
- Out-of-Distribution Detection with Deep Nearest NeighborsYiyou Sun, Yifei Ming, Xiaojin Zhu, Yixuan LiICML 2022 · 789 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language ModelsLonghui Yu, Weisen Jiang, Han Shi, Jincheng Yu et al.ICLR 2024 · 637 citations
Related papers
- Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language ModelsZizhuo Zhang, Jianing Zhu, Xinmu Ge, Zihua Zhao et al.ICLR 2026 · 16 citations
- Self-Consistency Preference OptimizationArchiki Prasad, Weizhe Yuan, Richard Yuanzhe Pang, Jing Xu et al.ICML 2025
- CREAM: Consistency Regularized Self-Rewarding Language ModelsZhaoyang Wang, Weilei He, Zhiyuan Liang, Xuchao Zhang et al.ICLR 2025
- Aligning Large Language Models via Fully Self-Synthetic DataShangjian Yin, Zhepei Wei, Xinyu Zhu, Wei-Lin Chen et al.ACL 2026 · 2 citations
- Large Language Models Can Self-ImproveJiaxin Huang, Shixiang Gu, Le Hou, Yuexin Wu et al.EMNLP 2023 · 184 citations
