Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following
Qingyu Ren, Qianyu He, Powei Chang, Jie Zeng, Zeye Sun, Fei Yu, Jiaqing Liang, Yanghua Xiao
摘要
Language models often struggle to follow multi-constraint instructions that are crucial for real-world applications. Existing reinforcement learning (RL) approaches suffer from dependency on external supervision and sparse reward signals from multi-constraint tasks. We propose a label-free self-supervised RL framework that eliminates dependency on external supervision by deriving reward signals directly from instructions and generating pseudo-labels for reward model training. Our approach introduces constraint decomposition strategies and efficient constraint-wise binary classification to address sparse reward challenges while maintaining computational efficiency. Experiments show that our approach generalizes well, achieving strong improvements across 3 indomain and 5 out-of-domain datasets, including challenging agentic and multi-turn instruction following. The code and data are publicly available at https://github.com/Rainier-rq/verlif and https://huggingface.co/dd12345789 . * Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Long Live The Balance: Information Bottleneck Driven Tree-based Policy OptimizationHao Jiang, Shurui Li, Tianpeng Bu, Bowen Xu 等ICML 2026
- PARIF: Pushing the Pareto Frontier of Instruction Following and Reasoning with Curriculum Reinforcement LearningRongchuan Mu, Zexin Wang, Qianyu Wang, Minghua Ma 等ACL 2026
- TinyJudge: Unverifiable Constraint Alignment via Lightweight Specialist EnsemblesYirong Zeng, Yufei Liu, Xiao Ding, Yutai Hou 等ACL 2026
它引用的顶会 Paper12
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li 等ICML 2024 · 被引用 569 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- COLLIE: Systematic Construction of Constrained Text Generation TasksShunyu Yao, Howard Chen, Austin W. Hanjie, Runzhe Yang 等ICLR 2024 · 被引用 65 次
- CFBench: A Comprehensive Constraints-Following Benchmark for LLMsTao Zhang, Chenglin Zhu, Yanjun Shen, Wenjing Luo 等ACL 2025 · 被引用 53 次
- VerIF: Verification Engineering for Reinforcement Learning in Instruction FollowingHao Peng, Yunjia Qi, Xiaozhi Wang, Bin Xu 等EMNLP 2025 · 被引用 24 次
相关 Paper
- ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language ModelsZirui Song, Guangxian Ouyang, Mingzhe Li, Yuheng Ji 等AAAI 2026 · 被引用 21 次
- DecIF: Improving Instruction-Following through DecompositionTingfeng Hui, Pengyu Zhu, Bowen Ping, Ling Tang 等ACL 2026
- Incentivizing Reasoning for Advanced Instruction-Following of Large Language ModelsYulei Qin, Gang Li, Zongyi Li, Zihan Xu 等NeurIPS 2025 · 被引用 17 次
- Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language ModelsZizhuo Zhang, Jianing Zhu, Xinmu Ge, Zihua Zhao 等ICLR 2026 · 被引用 16 次
- Supervised Reinforcement Learning: From Expert Trajectories to Step-wise ReasoningYihe Deng, I-Hung Hsu, Jun Yan, Zifeng Wang 等ICLR 2026 · 被引用 11 次
