Stackelberg Self-Annotation: A Robust Approach to Data-Efficient LLM Alignment
Chu Xu, Zhixin Zhang, Tianyu Jia, Yujie Jin
摘要
Aligning large language models (LLMs) with human preferences typically demands vast amounts of meticulously curated data, which is both expensive and prone to labeling noise. We propose Stackelberg Game Preference Optimization (SGPO), a robust alignment framework that models alignment as a two-player Stackelberg game between a policy (leader) and a worst-case preference distribution (follower). The proposed SGPO guarantees -bounded regret within an -Wasserstein ball, offering formal robustness to (self-)annotation noise. We instantiate SGPO with Stackelberg Self-Annotated Preference Optimization (SSAPO), which uses minimal human-labeled"seed"preferences and iteratively self-annotates new prompts. In each iteration, SSAPO applies a distributionally robust reweighting of synthetic annotations, ensuring that noisy or biased self-labels do not derail training. Remarkably, using only 2K seed preferences -- about 1/30 of standard human labels -- SSAPO achieves strong win rates against GPT-4 across multiple benchmarks within three iterations. These results highlight that a principled Stackelberg formulation yields data-efficient alignment for LLMs, significantly reducing reliance on costly human annotations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game PerspectiveHaichuan Wang, Tao Lin, Lingkai Kong, Ce Li 等ICML 2026 · 被引用 3 次
- AnomSeer: Reinforcing Multimodal LLMs to Reason for Time-Series Anomaly DetectionJunru Zhang, Lang Feng, Haoran Shi, Xu Guo 等ICML 2026
它引用的顶会 Paper18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky 等ICML 2024 · 被引用 973 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI FeedbackHarrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard 等ICML 2024 · 被引用 598 次
相关 Paper
- Reliability-Aware LLM Alignment from Inconsistent Human FeedbackJingyi Huang, Ruohan Zong, Yujun Feng, Liran Ma 等ICML 2026
- Self-Play Preference Optimization for Language Model AlignmentYue Wu, Zhiqing Sun, Huizhuo Yuan, Kaixuan Ji 等ICLR 2025
- RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM AlignmentXiaoyang Cao, Zelai Xu, Mo Guang, Kaiwen Long 等ICLR 2026 · 被引用 4 次
- Robust LLM Alignment via Distributionally Robust Direct Preference OptimizationZaiyan Xu, Sushil Vemuri, Kishan Panaganti, Dileep Kalathil 等NeurIPS 2025 · 被引用 18 次
- Aligning Large Language Models via Fully Self-Synthetic DataShangjian Yin, Zhepei Wei, Xinyu Zhu, Wei-Lin Chen 等ACL 2026 · 被引用 2 次
