Following Length Constraints in Instructions
Weizhe Yuan, Ilia Kulikov, Ping Yu, Kyunghyun Cho, Sainbayar Sukhbaatar, Jason E. Weston, Jing Xu
2025年份
1被引次数
20顶会引用
摘要
Aligned instruction following models can better fulfill user requests than their unaligned counterparts. However, it has been shown that there is a length bias in evaluation of such models, and that training algorithms tend to exploit this bias by learning longer responses. In this work we show how to train models that can be controlled at inference time with instructions containing desired length constraints. Such models are superior in length instructed evaluations, outperforming standard instruction following models such as GPT4, Llama 3 and Mixtral.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Scalable Chain of Thoughts via Elastic ReasoningYuhui Xu, Hanze Dong, Lei Wang, Doyen Sahoo 等ICLR 2026 · 被引用 42 次
- On Reasoning Strength Planning in Large Reasoning ModelsLeheng Sheng, An Zhang, Zijian Wu, Weixiang Zhao 等NeurIPS 2025 · 被引用 17 次
- Efficient Reasoning with Balanced ThinkingYulin Li, Tengyao Tu, Li Ding, Junjie Wang 等ICLR 2026 · 被引用 7 次
- Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-JudgeTianhao Wu, Weizhe Yuan, Olga Golovneva, Jing Xu 等EMNLP 2025 · 被引用 6 次
- Disentangling Length Bias in Preference Learning via Response-Conditioned ModelingJianfeng Cai, Jinhua Zhu, Ruopei Sun, Yue Wang 等ICLR 2026 · 被引用 6 次
它引用的顶会 Paper3
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
- Self-Alignment with Instruction BacktranslationXian Li, Ping Yu, Chunting Zhou, Timo Schick 等ICLR 2024 · 被引用 174 次
相关 Paper
- Long Is More for Alignment: A Simple but Tough-to-Beat Baseline for Instruction Fine-TuningHao Zhao, Maksym Andriushchenko, Francesco Croce, Nicolas FlammarionICML 2024 · 被引用 96 次
- LARFT: Closing the Cognition-Action Gap for Length Instruction Following in Large Language ModelsWei Zhang, Lintong Du, yuanhe zhang, Zhenhong Zhou 等ICML 2026
- Instruction Tuning With Loss Over InstructionsZhengxiang Shi, Adam X. Yang, Bin Wu, Laurence Aitchison 等NeurIPS 2024 · 被引用 55 次
- Length Controlled Generation for Black-box LLMsYuxuan Gu, Wenjie Wang, Xiaocheng Feng, Weihong Zhong 等ACL 2025
- UltraIF: Advancing Instruction Following from the WildKaikai An, Li Sheng, Ganqu Cui, Shuzheng Si 等EMNLP 2025 · 被引用 1 次
