TaskLAMA: Probing the Complex Task Understanding of Language Models
Quan Yuan, Mehran Kazemi, Xin Xu, Isaac Noble, Vaiva Imbrasaite, Deepak Ramachandran
Abstract
Structured Complex Task Decomposition (SCTD) is the problem of breaking down a complex real-world task (such as planning a wedding) into a directed acyclic graph over individual steps that contribute to achieving the task, with edges specifying temporal dependencies between them. SCTD is an important component of assistive planning tools, and a challenge for commonsense reasoning systems. We probe how accurately SCTD can be done with the knowledge extracted from Large Language Models (LLMs). We introduce a highquality human-annotated dataset for this problem and novel metrics to fairly assess performance of LLMs against several baselines. Our experiments reveal that LLMs are able to decompose complex tasks into individual steps effectively, with a relative improvement of 15% to 280% over the best baseline. We also propose a number of approaches to further improve their performance, with a relative improvement of 7% to 37% over the base model. However, we find that LLMs still struggle to predict pairwise temporal dependencies, which reveals a gap in their understanding of complex tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d1b6ac13-c9f6-48da-b184-447c2f3e0162Cited by top-tier papers6
- PlanGenLLMs: A Modern Survey of LLM Planning CapabilitiesHui Wei, Zihao Zhang, Shenghua He, Tian Xia et al.ACL 2025 · 78 citations
- Progress Reward Model for Reinforcement Learning via Large Language ModelsXiuhui Zhang, Ning Gao, Xingyu Jiang, Yihui Chen et al.NeurIPS 2025 · 3 citations
- Open Grounded Planning: Challenges and Benchmark ConstructionShiguang Guo, Ziliang Deng, Hongyu Lin, Yaojie Lu et al.ACL 2024 · 2 citations
- Log2Plan: An Adaptive GUI Automation Framework Integrated with Task Mining ApproachSeoyoung Lee, Seobin Yoon, Seongbeen Lee, Hyesoo Kim et al.UIST 2025 · 2 citations
- Multi-modal Sketch-Based Behavior Tree SynthesisWenmeng Zhang, Zhenbang Chen, Weijiang HongOOPSLA 2025
Builds on7
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- Evaluating Commonsense in Pre-Trained Language ModelsXuhui Zhou, Yue Zhang, Leyang Cui, Dandan HuangAAAI 2020 · 198 citations
- Language Models of Code are Few-Shot Commonsense LearnersAman Madaan, Shuyan Zhou, Uri Alon, Yiming Yang et al.EMNLP 2022 · 103 citations
- The Power of Scale for Parameter-Efficient Prompt TuningBrian Lester, Rami Al-Rfou, Noah ConstantEMNLP 2021 · 94 citations
Related papers
- Large Language Models as Commonsense Knowledge for Large-Scale Task PlanningZirui Zhao, Wee Sun Lee, David HsuNeurIPS 2023 · 423 citations
- PAGED: A Benchmark for Procedural Graphs Extraction from DocumentsWeihong Du, Wenrui Liao, Hongru Liang, Wenqiang LeiACL 2024 · 4 citations
- Neuro-Symbolic Procedural Planning with Commonsense PromptingYujie Lu, Weixi Feng, Wanrong Zhu, Wenda Xu et al.ICLR 2023 · 3 citations
- SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional VideosYulei Niu, Wenliang Guo, Long Chen, Xudong Lin et al.ICLR 2024 · 26 citations
- Multi-Modal Grounded Planning and Efficient Replanning for Learning Embodied Agents with a Few ExamplesTaewoong Kim, Byeonghwi Kim, Jonghyun ChoiAAAI 2025 · 8 citations
