Lune

ACL2026顶会

CoAct: Co-Active LLM Preference Learning with Human-AI Synergy

Ruiyao Xu, Mihir Parmar, Tiankai Yang, Zhengyu Hu, Yue Zhao, Kaize Ding

2026年份
1被引次数

摘要

Learning from preference-based feedback has become an effective approach for aligning LLMs across diverse tasks. However, highquality human-annotated preference data remains expensive and scarce. Existing methods address this challenge through either selfrewarding, which scales by using purely AIgenerated labels but risks unreliability, or active learning, which ensures quality through oracle annotation but cannot fully leverage unlabeled data. In this paper, we present COACT, a novel framework that synergistically combines selfrewarding and active learning through strategic human-AI collaboration. COACT leverages self-consistency to identify both reliable selflabeled data and samples that are requiring oracle verification. Additionally, oracle feedback guides the model to generate new instructions within its solvable capability. Evaluated on three reasoning benchmarks across two model families, COACT achieves average improvements of +13.25% on GSM8K, +8.19% on MATH, and +13.16% on WebInstruct, consistently outperforming all baselines. 1 1 Our code is available at https://github.com/rux001/ CoAct .

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper26

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖