ASPERA: A Simulated Environment to Evaluate Planning for Complex Action Execution
Alexandru Coca, Mark Gaynor, Zhenxing Zhang, Jianpeng Cheng, Bo-Hsiang Tseng, Peter Boothroyd, Héctor Martínez Alonso, Diarmuid Ó Séaghdha, Anders Johannsen
Abstract
This work evaluates the potential of large language models (LLMs) to power digital assistants capable of complex action execution. These assistants rely on pre-trained programming knowledge to execute multi-step goals by composing objects and functions defined in assistant libraries into action execution programs. To achieve this, we develop ASPERA, a framework comprising an assistant library simulation and a human-assisted LLM data generation engine. Our engine allows developers to guide LLM generation of high-quality tasks consisting of complex user queries, simulation state and corresponding validation programs, tackling data availability and evaluation robustness challenges. Alongside the framework we release Asper-Bench, an evaluation dataset of 250 challenging tasks generated using ASPERA, which we use to show that program generation grounded in custom assistant libraries is a significant challenge to LLMs compared to dependency-free code generation 1 . * Work done while at Apple. 1 https://github.com/apple/ml-aspera . 1 Output programs with cancellations 2 3 Execute "Cancel lunch with Jill" . . .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 313f0368-094f-45a7-b60a-e26699655872Builds on19
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 1,477 citations
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu et al.ICLR 2024 · 1,469 citations
Related papers
- LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model ProgramsYunsheng Ma, Can Cui, Xu Cao, Wenqian Ye et al.CVPR 2024 · 39 citations
- Towards Synthetic Trace Generation of Modeling Operations using In-Context Learning ApproachVittoriano Muttillo, Claudio Di Sipio, Riccardo Rubei, Luca Berardinelli et al.ASE 2024 · 1 citation
- XCodeEval: An Execution-based Large Scale Multilingual Multitask Benchmark for Code Understanding, Generation, Translation and RetrievalMohammad Abdullah Matin Khan, M. Saiful Bari, Xuan Do Long, Weishi Wang et al.ACL 2024 · 21 citations
- AppBench: Planning of Multiple APIs from Various APPs for Complex User InstructionHongru Wang, Rui Wang, Boyang Xue, Heming Xia et al.EMNLP 2024 · 2 citations
- ARBench: Algorithmic Reasoner or API Alchemist? Evaluating LLMs Beyond API CallsRenbiao Liu, Chao-Zeng Ma, Anqi Li, Hui Sun et al.AAAI 2026
