ActPlan-1K: Benchmarking the Procedural Planning Ability of Visual Language Models in Household Activities
Ying Su, Zhan Ling, Haochen Shi, Cheng Jiayang, Yauwai Yim, Yangqiu Song
摘要
Large language models (LLMs) have been adopted to process textual task description and accomplish procedural planning in embodied AI tasks because of their powerful reasoning ability. However, there is still lack of study on how vision language models (VLMs) behave when multi-modal task inputs are considered. Counterfactual planning that evaluates the model's reasoning ability over alternative task situations are also under exploited. In order to evaluate the planning ability of both multimodal and counterfactual aspects, we propose ActPlan-1K. ActPlan-1K is a multi-modal planning benchmark constructed based on ChatGPT and household activity simulator iGibson2. The benchmark consists of 153 activities and 1,187 instances. Each instance describing one activity has a natural language task description and multiple environment images from the simulator. The gold plan of each instance is action sequences over the objects in provided scenes. Both the correctness and commonsense satisfaction are evaluated on typical VLMs. It turns out that current VLMs are still struggling at generating human-level procedural plans for both normal activities and counterfactual activities. We further provide automatic evaluation metrics by finetuning over BLEURT model to facilitate future research on our benchmark. Generate admissible procedural plans for assembling gift baskets. There are four baskets on the floor, cookies, cheeses, and bows on the tables. The goal is to place one of each item inside on of the baskets as shown in the images. Gold Plan 1. walk_to(table) 2. grab(cookie) 3.walk_to(basket_1) 4. place_inside(basket_1, cookie) 5. walk_to(table) 6. grab(cheese) 7. walk_to(basket_1) 8. ……
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural PlanningShibo Sun, Xue Li, Donglin Di, Mingjie Wei 等ACM MM 2025 · 被引用 4 次
- Enhancing LLM Planning for Robotics Manipulation through Hierarchical Procedural Knowledge GraphsJiacong Zhou, Jiaxu Miao, Xianyun Wang, Jun YuNeurIPS 2025 · 被引用 1 次
- iVISPAR - An Interactive Visual-Spatial Reasoning Benchmark for VLMsJulius Mayer, Mohamad Ballout, Serwan Jassim, Farbod Nosrat Nezami 等EMNLP 2025 · 被引用 1 次
- VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical ReasoningNilay Yilmaz, Maitreya Patel, Yiran Lawrence Luo, Tejas Gokhale 等ICLR 2025
- PEAP: Proactive Embodied Action Sequence Planning with Joint Understanding of Vision and Audio PerceptionTianwei Lan, Jiaqi Wu, Zeming Liu, Zhaoxin Fan 等ACL 2026
它引用的顶会 Paper11
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 被引用 1,539 次
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao 等ICCV 2023 · 被引用 685 次
- Pre-Trained Language Models for Interactive Decision-MakingShuang Li, Xavier Puig, Chris Paxton, Yilun Du 等NeurIPS 2022 · 被引用 341 次
相关 Paper
- ACPBench: Reasoning About Action, Change, and PlanningHarsha Kokel, Michael Katz, Kavitha Srinivas, Shirin SohrabiAAAI 2025 · 被引用 35 次
- VirtualEnv: A Platform for Embodied AI ResearchKabir Swain, Sijie Han, Ayush Raina, Jin Zhang 等AAAI 2026
- ProcWorld: Benchmarking Large Model Planning in Reachability-Constrained EnvironmentsDong Wang, Xinghang Li, Zhengshen Zhang, Jirong Liu 等EMNLP 2025
- AmbiK: Dataset of Ambiguous Tasks in Kitchen EnvironmentAnastasiia Ivanova, Eva Bakaeva, Zoya Volovikova, Alexey K. Kovalev 等ACL 2025
- MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGIKaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li 等ICML 2024 · 被引用 184 次
