EIFBENCH: Extremely Complex Instruction Following Benchmark for Large Language Models
Tao Zou, Xinghua Zhang, Haiyang Yu, Minzheng Wang, Fei Huang, Yongbin Li
摘要
With the development and widespread application of large language models (LLMs), the new paradigm of "Model as Product" is rapidly evolving, and demands higher capabilities to address complex user needs, often requiring precise workflow execution which involves the accurate understanding of multiple tasks. However, existing benchmarks focusing on single-task environments with limited constraints lack the complexity required to fully reflect real-world scenarios. To bridge this gap, we present the Extremely Complex Instruction Following Benchmark (EIFBENCH), meticulously crafted to facilitate a more realistic and robust evaluation of LLMs. EIFBENCH not only includes multi-task scenarios that enable comprehensive assessment across diverse task types concurrently, but also integrates a variety of constraints, replicating complex operational environments. Furthermore, we propose the Segment Policy Optimization (SegPO) algorithm to enhance the LLM's ability to accurately fulfill multi-task workflow. Evaluations on EIFBENCH have unveiled considerable performance discrepancies in existing LLMs when challenged with these extremely complex instructions. This finding underscores the necessity for ongoing optimization to navigate the intricate challenges posed by LLM applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- RAIN-Merging: A Gradient-Free Method to Enhance Instruction Following in Large Reasoning Models with Preserved Thinking FormatZhehao Huang, Yuhang Liu, Baijiong Lin, Yixin Lou 等ICLR 2026 · 被引用 7 次
- One Battle After Another: Probing LLMs' Limits on Multi-Turn Instruction Following with a Benchmark Evolving FrameworkQi Jia, Ye Shen, Xiujie Song, Kaiwei Zhang 等ACL 2026 · 被引用 3 次
- Revisiting the Reliability of Language Models in Instruction-FollowingJianshuo Dong, Yutong Zhang, Liu Yan, Zhenyu Zhong 等ACL 2026 · 被引用 3 次
- ImpRIF: Stronger Implicit Reasoning Leads to Better Complex Instruction FollowingYuancheng Yang, Lin Yang, Xu Wang, Chao Tong 等ACL 2026 · 被引用 1 次
- Implicit Intelligence - Evaluating Agents on What Users Don’t SayVed Sirdeshmukh, Marc WetterICML 2026
它引用的顶会 Paper13
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
- Cross-Task Generalization via Natural Language Crowdsourcing InstructionsSwaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh HajishirziACL 2022 · 被引用 887 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- Evaluating Large Language Models at Evaluating Instruction FollowingZhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng 等ICLR 2024 · 被引用 299 次
相关 Paper
- Can Large Language Models Understand Real-World Complex Instructions?Qianyu He, Jie Zeng, Wenhao Huang, Lina Chen 等AAAI 2024 · 被引用 99 次
- CFBench: A Comprehensive Constraints-Following Benchmark for LLMsTao Zhang, Chenglin Zhu, Yanjun Shen, Wenjing Luo 等ACL 2025 · 被引用 53 次
- Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?Guijin Son, Sangwon Baek, Sangdae Nam, Ilgyun Jeong 等ACL 2024
- OrchestrationBench: LLM-Driven Agentic Planning and Tool Use in Multi-Domain ScenariosAelim Ahn, Sooyeon Lee, Hyosun Wang, Chiwan Park 等ICLR 2026
- MaXIFE: Multilingual and Cross-lingual Instruction Following EvaluationYile Liu, Ziwei Ma, Xiu Jiang, Jinglu Hu 等ACL 2025 · 被引用 5 次
