Lune

ACL2026Top-tier venue

NL Schedule: Evaluate Multitask Scheduling Capability of Large Language Models

Wenrui Liao, Weihong Du, Yi Li, Hongru Liang, Wenqiang Lei

2026Year

Abstract

Automated schedule generation for multitask from natural language descriptions has huge potential in modern industry. While classic methods bypass language complexities by using preformatted matrices, and recent LLM+solver approaches introduce new fragilities by relying on solver-specific code generation. This raises critical questions: Can large language models (LLMs) solve this NL ⇒ Schedule task end-to-end well (RQ1)? If the answer is "no", where do they fall short (RQ2)? And how can their capabilities be enhanced (RQ3)? To answer these questions, we introduce NL ⇒ Schedule, the first benchmark for this task, equipped with a dataset of 240 descriptionschedule pairs constructed from real-world materials and a rigorous evaluation suite. Our evaluation of nine state-of-the-art LLMs reveals the limitations of different LLMs in procedure grounding and the strengths of advanced LLMs in global planning via local analysis. To address these shortcomings, we propose MANS, a novel multi-agent framework. Extensive experiments show that MANS achieves more robust performance comparable to six state-ofthe-art LLM+solver methods. We hope NL ⇒ Schedule and MANS will serve as a solid foundation for automatic scheduling. The code and dataset are available in https://github.com/ SCUNLP/NL2Schedule

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 958167f2-34e5-405d-a0d6-56e1a62323cb

Builds on5

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines