FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models
Yuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong, Liangyou Li, Fei Mi, Lifeng Shang, Xin Jiang, Qun Liu, Wei Wang
Abstract
The ability to follow instructions is crucial for Large Language Models (LLMs) to handle various real-world applications. Existing benchmarks primarily focus on evaluating pure response quality, rather than assessing whether the response follows constraints stated in the instruction. To fill this research gap, in this paper, we propose FollowBench, a Multi-level Fine-grained Constraints Following Benchmark for LLMs. FollowBench comprehensively includes five different types (i.e., Content, Situation, Style, Format, and Example) of fine-grained constraints. To enable a precise constraint following estimation on diverse difficulties, we introduce a Multi-level mechanism that incrementally adds a single constraint to the initial instruction at each increased level. To assess whether LLMs' outputs have satisfied every individual constraint, we propose to prompt strong LLMs with constraint-evolution paths to handle challenging open-ended instructions. By evaluating 13 closed-source and opensource popular LLMs on FollowBench, we highlight the weaknesses of LLMs in instruction following and point towards potential avenues for future work. The data and code are publicly available at https://github. com/YJiangcm/FollowBench .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7a76e217-b822-4c1d-a5b8-ed0d3fc1ace2Cited by top-tier papers64
- Checklists Are Better Than Reward Models For Aligning Language ModelsVijay Viswanathan, Yanchao Sun, Xiang Kong, Meng Cao et al.NeurIPS 2025 · 127 citations
- OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM AlignmentTianci Liu, Ran Xu, Tony Yu, Ilgee Hong et al.ACL 2026 · 75 citations
- CFBench: A Comprehensive Constraints-Following Benchmark for LLMsTao Zhang, Chenglin Zhu, Yanjun Shen, Wenjing Luo et al.ACL 2025 · 53 citations
- AdvancedIF: Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction FollowingYun He, Wenzhe Li, Hejia Zhang, Songlin Li et al.ACL 2026 · 36 citations
- IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference OptimizationXinghua Zhang, Haiyang Yu, Cheng Fu, Fei Huang et al.ACL 2025 · 29 citations
Builds on11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- Cross-Task Generalization via Natural Language Crowdsourcing InstructionsSwaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh HajishirziACL 2022 · 887 citations
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
Related papers
- SysBench: Can LLMs Follow System Message?Yanzhao Qin, Tao Zhang, Tao Zhang, Yanjun Shen et al.ICLR 2025
- Can Large Language Models Understand Real-World Complex Instructions?Qianyu He, Jie Zeng, Wenhao Huang, Lina Chen et al.AAAI 2024 · 99 citations
- ConInstruct: Evaluating Large Language Models on Conflict Detection and Resolution in InstructionsXingwei He, Qianru Zhang, Pengfei Chen, Guanhua Chen et al.AAAI 2026 · 2 citations
- LongGenBench: Benchmarking Long-Form Generation in Long Context LLMsYuhao Wu, Ming Shan Hee, Zhiqiang Hu, Roy Ka-Wei LeeICLR 2025
- IF-RewardBench: Benchmarking Judge Models for Instruction-Following EvaluationBosi Wen, Yilin Niu, Cunxiang Wang, Xiaoying Ling et al.ACL 2026 · 3 citations
