Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following
Qingyu Ren, Qianyu He, Powei Chang, Jie Zeng, Zeye Sun, Fei Yu, Jiaqing Liang, Yanghua Xiao
Abstract
Language models often struggle to follow multi-constraint instructions that are crucial for real-world applications. Existing reinforcement learning (RL) approaches suffer from dependency on external supervision and sparse reward signals from multi-constraint tasks. We propose a label-free self-supervised RL framework that eliminates dependency on external supervision by deriving reward signals directly from instructions and generating pseudo-labels for reward model training. Our approach introduces constraint decomposition strategies and efficient constraint-wise binary classification to address sparse reward challenges while maintaining computational efficiency. Experiments show that our approach generalizes well, achieving strong improvements across 3 indomain and 5 out-of-domain datasets, including challenging agentic and multi-turn instruction following. The code and data are publicly available at https://github.com/Rainier-rq/verlif and https://huggingface.co/dd12345789 . * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ee50293-9309-4d47-9be9-1d03d08f0d5fCited by top-tier papers3
- Long Live The Balance: Information Bottleneck Driven Tree-based Policy OptimizationHao Jiang, Shurui Li, Tianpeng Bu, Bowen Xu et al.ICML 2026
- PARIF: Pushing the Pareto Frontier of Instruction Following and Reasoning with Curriculum Reinforcement LearningRongchuan Mu, Zexin Wang, Qianyu Wang, Minghua Ma et al.ACL 2026
- TinyJudge: Unverifiable Constraint Alignment via Lightweight Specialist EnsemblesYirong Zeng, Yufei Liu, Xiao Ding, Yutai Hou et al.ACL 2026
Builds on12
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li et al.ICML 2024 · 569 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- COLLIE: Systematic Construction of Constrained Text Generation TasksShunyu Yao, Howard Chen, Austin W. Hanjie, Runzhe Yang et al.ICLR 2024 · 65 citations
- CFBench: A Comprehensive Constraints-Following Benchmark for LLMsTao Zhang, Chenglin Zhu, Yanjun Shen, Wenjing Luo et al.ACL 2025 · 53 citations
- VerIF: Verification Engineering for Reinforcement Learning in Instruction FollowingHao Peng, Yunjia Qi, Xiaozhi Wang, Bin Xu et al.EMNLP 2025 · 24 citations
Related papers
- ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language ModelsZirui Song, Guangxian Ouyang, Mingzhe Li, Yuheng Ji et al.AAAI 2026 · 21 citations
- DecIF: Improving Instruction-Following through DecompositionTingfeng Hui, Pengyu Zhu, Bowen Ping, Ling Tang et al.ACL 2026
- Incentivizing Reasoning for Advanced Instruction-Following of Large Language ModelsYulei Qin, Gang Li, Zongyi Li, Zihan Xu et al.NeurIPS 2025 · 17 citations
- Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language ModelsZizhuo Zhang, Jianing Zhu, Xinmu Ge, Zihua Zhao et al.ICLR 2026 · 16 citations
- Supervised Reinforcement Learning: From Expert Trajectories to Step-wise ReasoningYihe Deng, I-Hung Hsu, Jun Yan, Zifeng Wang et al.ICLR 2026 · 11 citations
