Following Length Constraints in Instructions
Weizhe Yuan, Ilia Kulikov, Ping Yu, Kyunghyun Cho, Sainbayar Sukhbaatar, Jason E. Weston, Jing Xu
Abstract
Aligned instruction following models can better fulfill user requests than their unaligned counterparts. However, it has been shown that there is a length bias in evaluation of such models, and that training algorithms tend to exploit this bias by learning longer responses. In this work we show how to train models that can be controlled at inference time with instructions containing desired length constraints. Such models are superior in length instructed evaluations, outperforming standard instruction following models such as GPT4, Llama 3 and Mixtral.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- Scalable Chain of Thoughts via Elastic ReasoningYuhui Xu, Hanze Dong, Lei Wang, Doyen Sahoo et al.ICLR 2026 · 42 citations
- On Reasoning Strength Planning in Large Reasoning ModelsLeheng Sheng, An Zhang, Zijian Wu, Weixiang Zhao et al.NeurIPS 2025 · 17 citations
- Efficient Reasoning with Balanced ThinkingYulin Li, Tengyao Tu, Li Ding, Junjie Wang et al.ICLR 2026 · 7 citations
- Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-JudgeTianhao Wu, Weizhe Yuan, Olga Golovneva, Jing Xu et al.EMNLP 2025 · 6 citations
- Disentangling Length Bias in Preference Learning via Response-Conditioned ModelingJianfeng Cai, Jinhua Zhu, Ruopei Sun, Yue Wang et al.ICLR 2026 · 6 citations
Builds on3
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang et al.NeurIPS 2023 · 948 citations
- Self-Alignment with Instruction BacktranslationXian Li, Ping Yu, Chunting Zhou, Timo Schick et al.ICLR 2024 · 174 citations
Related papers
- Long Is More for Alignment: A Simple but Tough-to-Beat Baseline for Instruction Fine-TuningHao Zhao, Maksym Andriushchenko, Francesco Croce, Nicolas FlammarionICML 2024 · 96 citations
- LARFT: Closing the Cognition-Action Gap for Length Instruction Following in Large Language ModelsWei Zhang, Lintong Du, yuanhe zhang, Zhenhong Zhou et al.ICML 2026
- Instruction Tuning With Loss Over InstructionsZhengxiang Shi, Adam X. Yang, Bin Wu, Laurence Aitchison et al.NeurIPS 2024 · 55 citations
- Length Controlled Generation for Black-box LLMsYuxuan Gu, Wenjie Wang, Xiaocheng Feng, Weihong Zhong et al.ACL 2025
- UltraIF: Advancing Instruction Following from the WildKaikai An, Li Sheng, Ganqu Cui, Shuzheng Si et al.EMNLP 2025 · 1 citation
