Lune

ACL2026Top-tier venue

CoAct: Co-Active LLM Preference Learning with Human-AI Synergy

Ruiyao Xu, Mihir Parmar, Tiankai Yang, Zhengyu Hu, Yue Zhao, Kaize Ding

2026Year
1Citations

Abstract

Learning from preference-based feedback has become an effective approach for aligning LLMs across diverse tasks. However, highquality human-annotated preference data remains expensive and scarce. Existing methods address this challenge through either selfrewarding, which scales by using purely AIgenerated labels but risks unreliability, or active learning, which ensures quality through oracle annotation but cannot fully leverage unlabeled data. In this paper, we present COACT, a novel framework that synergistically combines selfrewarding and active learning through strategic human-AI collaboration. COACT leverages self-consistency to identify both reliable selflabeled data and samples that are requiring oracle verification. Additionally, oracle feedback guides the model to generate new instructions within its solvable capability. Evaluated on three reasoning benchmarks across two model families, COACT achieves average improvements of +13.25% on GSM8K, +8.19% on MATH, and +13.16% on WebInstruct, consistently outperforming all baselines. 1 1 Our code is available at https://github.com/rux001/ CoAct .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 0a5506f0-7b29-4047-95f4-d6c790e63f0e

Builds on26

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines