From Sub-Ability Diagnosis to Human-Aligned Generation: Bridging the Gap for Text Length Control via MarkerGen
Peiwen Yuan, Chuyi Tan, Shaoxiong Feng, Yiwei Li, Xinglin Wang, Yueqi Zhang, Jiayi Shi, Boyuan Pan, Yao Hu, Kan Li
Abstract
Despite the rapid progress of large language models (LLMs), their length-controllable text generation (LCTG) ability remains below expectations, posing a major limitation for practical applications. Existing methods mainly focus on end-to-end training to reinforce adherence to length constraints. However, the lack of decomposition and targeted enhancement of LCTG sub-abilities restricts further progress. To bridge this gap, we conduct a bottom-up decomposition of LCTG sub-abilities with human patterns as reference and perform a detailed error analysis. On this basis, we propose MARK-ERGEN, a simple-yet-effective plug-and-play approach that: (1) mitigates LLM fundamental deficiencies via external tool integration; (2) conducts explicit length modeling with dynamically inserted markers; (3) employs a threestage generation scheme to better align length constraints while maintaining content quality. Comprehensive experiments demonstrate that MARKERGEN significantly improves LCTG across various settings, exhibiting outstanding effectiveness and generalizability. 1 . * Equal contribution. † Corresponding author. 1 Our code have been released on https://github.com/ chuyi369/MarkerGen .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ea1d5ec0-a797-4c8f-983a-15b2fd7b1ac7Cited by top-tier papers1
Ask how each one uses itBuilds on7
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
- LLM-Powered Benchmark Factory: Reliable, Generic, and EfficientPeiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang et al.ACL 2026 · 8 citations
- Hansel: Output Length Controlling Framework for Large Language ModelsSeoha Song, Junhyun Lee, Hyeonmok KoAAAI 2025 · 2 citations
- Following Length Constraints in InstructionsWeizhe Yuan, Ilia Kulikov, Ping Yu, Kyunghyun Cho et al.EMNLP 2025 · 1 citation
Related papers
- Control Large Language Models via Divide and ConquerBingxuan Li, Yiwei Wang, Tao Meng, Kai-Wei Chang et al.EMNLP 2024 · 1 citation
- ToolGen: Unified Tool Retrieval and Calling via GenerationRenxi Wang, Xudong Han, Lei Ji, Shu Wang et al.ICLR 2025
- Fine-Grained Controllable Text Generation Using Non-Residual PromptingFredrik Carlsson, Joey Öhman, Fangyu Liu, Severine Verlinden et al.ACL 2022
- Adaptable Logical Control for Large Language ModelsHonghua Zhang, Po-Nien Kung, Masahiro Yoshida, Guy Van den Broeck et al.NeurIPS 2024 · 42 citations
- MetaAgent: Automatically Constructing Multi-Agent Systems Based on Finite State MachinesYaolun Zhang, Xiaogeng Liu, Chaowei XiaoICML 2025
