NL Schedule: Evaluate Multitask Scheduling Capability of Large Language Models
Wenrui Liao, Weihong Du, Yi Li, Hongru Liang, Wenqiang Lei
摘要
Automated schedule generation for multitask from natural language descriptions has huge potential in modern industry. While classic methods bypass language complexities by using preformatted matrices, and recent LLM+solver approaches introduce new fragilities by relying on solver-specific code generation. This raises critical questions: Can large language models (LLMs) solve this NL ⇒ Schedule task end-to-end well (RQ1)? If the answer is "no", where do they fall short (RQ2)? And how can their capabilities be enhanced (RQ3)? To answer these questions, we introduce NL ⇒ Schedule, the first benchmark for this task, equipped with a dataset of 240 descriptionschedule pairs constructed from real-world materials and a rigorous evaluation suite. Our evaluation of nine state-of-the-art LLMs reveals the limitations of different LLMs in procedure grounding and the strengths of advanced LLMs in global planning via local analysis. To address these shortcomings, we propose MANS, a novel multi-agent framework. Extensive experiments show that MANS achieves more robust performance comparable to six state-ofthe-art LLM+solver methods. We hope NL ⇒ Schedule and MANS will serve as a solid foundation for automatic scheduling. The code and dataset are available in https://github.com/ SCUNLP/NL2Schedule
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- On the Planning Abilities of Large Language Models - A Critical InvestigationKarthik Valmeekam, Matthew Marquez, Sarath Sreedharan, Subbarao KambhampatiNeurIPS 2023 · 被引用 509 次
- Learning to Dispatch for Job Shop Scheduling via Deep Reinforcement LearningCong Zhang, Wen Song, Zhiguang Cao, Jie Zhang 等NeurIPS 2020 · 被引用 497 次
- OptiMUS: Scalable Optimization Modeling with (MI)LP Solvers and Large Language ModelsAli AhmadiTeshnizi, Wenzhi Gao, Madeleine UdellICML 2024 · 被引用 77 次
- Re-examining the Role of Schema Linking in Text-to-SQLWenqiang Lei, Weixin Wang, Zhixin Ma, Tian Gan 等EMNLP 2020 · 被引用 71 次
- PAGED: A Benchmark for Procedural Graphs Extraction from DocumentsWeihong Du, Wenrui Liao, Hongru Liang, Wenqiang LeiACL 2024 · 被引用 4 次
相关 Paper
- PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent TasksMatthew Chang, Gunjan Chhablani, Alexander Clegg, Mikael Dallaire Cote 等ICLR 2025
- PrepBench: How Far Are We from Natural-Language-Driven Data Preparation?Jingzhe Xu, Rui Wang, Jiannan Wang, Guoliang LiVLDB 2026 · 被引用 3 次
- SheetCopilot: Bringing Software Productivity to the Next Level through Large Language ModelsHongxin Li, Jingran Su, Yuntao Chen, Qing Li 等NeurIPS 2023 · 被引用 75 次
- LoTa-Bench: Benchmarking Language-oriented Task Planners for Embodied AgentsJae-Woo Choi, Youngwoo Yoon, Hyobin Ong, Jaehong Kim 等ICLR 2024 · 被引用 49 次
- When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task DescriptionsMaya Larbi, Amal Akli, Mike Papadakis, Rihab Bouyousfi 等ICSE 2026
