Can Large Language Models Understand Real-World Complex Instructions?
Qianyu He, Jie Zeng, Wenhao Huang, Lina Chen, Jin Xiao, Qianxi He, Xunzhe Zhou, Jiaqing Liang, Yanghua Xiao
摘要
Large language models (LLMs) can understand human instructions, showing their potential for pragmatic applications beyond traditional NLP tasks. However, they still struggle with complex instructions, which can be either complex task descriptions that require multiple tasks and constraints, or complex input that contains long context, noise, heterogeneous information and multi-turn format. Due to these features, LLMs often ignore semantic constraints from task descriptions, generate incorrect formats, violate length or sample count constraints, and be unfaithful to the input text. Existing benchmarks are insufficient to assess LLMs' ability to understand complex instructions, as they are close-ended and simple. To bridge this gap, we propose CELLO, a benchmark for evaluating LLMs' ability to follow complex instructions systematically. We design eight features for complex instructions and construct a comprehensive evaluation dataset from real-world scenarios. We also establish four criteria and develop corresponding metrics, as current ones are inadequate, biased or too strict and coarse-grained. We compare the performance of representative Chinese-oriented and English-oriented models in following complex instructions through extensive experiments. Resources of CELLO are publicly available at https://github.com/Abbey4799/CELLO .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?Zhanke Zhou, Rong Tao, Jianing Zhu, Yiwen Luo 等NeurIPS 2024 · 被引用 74 次
- Fake News in Sheep's Clothing: Robust Fake News Detection Against LLM-Empowered Style AttacksJiaying Wu, Jiafeng Guo, Bryan HooiKDD 2024 · 被引用 69 次
- Fine-Tuned Large Language Model for Visualization System: A Study on Self-Regulated Learning in EducationLin Gao, Jing Lu, Zekai Shao, Ziyue Lin 等IEEE VIS 2024 · 被引用 27 次
- LIFBench: Evaluating the Instruction Following Performance and Stability of Large Language Models in Long-Context ScenariosXiaodong Wu, Minhao Wang, Yichen Liu, Xiaoming Shi 等ACL 2025 · 被引用 22 次
- Incentivizing Reasoning for Advanced Instruction-Following of Large Language ModelsYulei Qin, Gang Li, Zongyi Li, Zihan Xu 等NeurIPS 2025 · 被引用 17 次
它引用的顶会 Paper10
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu 等ICLR 2024 · 被引用 1,469 次
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun 等ICLR 2024 · 被引用 945 次
- GLM-130B: An Open Bilingual Pre-trained ModelAohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang 等ICLR 2023 · 被引用 295 次
- Controlled Text Generation with Natural Language InstructionsWangchunshu Zhou, Yuchen Eleanor Jiang, Ethan Wilcox, Ryan Cotterell 等ICML 2023 · 被引用 121 次
- Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat DataCanwen Xu, Daya Guo, Nan Duan, Julian J. McAuleyEMNLP 2023 · 被引用 112 次
相关 Paper
- CFBench: A Comprehensive Constraints-Following Benchmark for LLMsTao Zhang, Chenglin Zhu, Yanjun Shen, Wenjing Luo 等ACL 2025 · 被引用 53 次
- FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language ModelsYuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong 等ACL 2024 · 被引用 10 次
- LongGenBench: Benchmarking Long-Form Generation in Long Context LLMsYuhao Wu, Ming Shan Hee, Zhiqiang Hu, Roy Ka-Wei LeeICLR 2025
- LexInstructEval: Lexical Instruction Following Evaluation for Large Language ModelsHuimin Ren, Yan Liang, Baiqiao Su, Chaobo Sun 等AAAI 2026
- EIFBENCH: Extremely Complex Instruction Following Benchmark for Large Language ModelsTao Zou, Xinghua Zhang, Haiyang Yu, Minzheng Wang 等EMNLP 2025 · 被引用 1 次
