Lune

ICML2026顶会

LARFT: Closing the Cognition-Action Gap for Length Instruction Following in Large Language Models

Wei Zhang, Lintong Du, yuanhe zhang, Zhenhong Zhou, Kun Wang, Li Sun, Sen Su

2026年份

摘要

Despite the strong performance of Large Language Models (LLMs) on complex instruction-following tasks, precise control of output length remains a persistent challenge. Existing methods primarily attempt to enforce length constraints by externally imposing length signals or optimization objectives, while largely overlooking the underlying limitation: the model's intrinsic deficit in length cognition. To address this, we propose LARFT ( L ength- A ware R einforcement F ine- T uning), a training framework that aligns the model's length cognition with its action. Specifically, LARFT integrates length-oriented reinforcement learning with a hindsight length awareness. By transforming on-policy data into hindsight self-awareness tasks where the model learns to identify the actual length of its own generation, LARFT jointly optimizes the model’s internal representation of length information and refines its policy to satisfy length constraints, thereby achieving precise and reliable length instruction following. Extensive experiments across four base models demonstrate that LARFT outperforms existing baselines, achieving an average improvement of +20.92 points across three length instruction following benchmarks with only a marginal decline of -1.45 points on four general capability benchmarks. Our code is available at https://github.com/Captain-zhangw/LARFT.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 00b3ffa0-16fe-408b-8141-4f92f34bcc5e

它引用的顶会 Paper12

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖