PEAP: Proactive Embodied Action Sequence Planning with Joint Understanding of Vision and Audio Perception
Tianwei Lan, Jiaqi Wu, Zeming Liu, Zhaoxin Fan, Haifeng Wang, Yuhang Guo
摘要
Embodied Action Sequence Planning focuses on the capability of embodied agents to implement action planning via environmental perception. This technology enables diverse intelligent assistance for real-world scenarios such as home and office environments. To address the limitations of existing embodied agents in meeting the requirement for proactivity and achieving joint understanding of visual and audio information, this study investigates the ability of embodied agents to proactively provide assistance through action sequence planning based on joint understanding of vision and audio perception without explicit human instructions. Correspondingly, we propose PEAP, the first multimodal proactive embodied action sequence planning dataset. We evaluate the performance of multiple Large Language Models on the PEAP dataset. The results demonstrate that these models still exhibit significant deficiencies on this task particularly lacking accurate environmental perception capabilities. Furthermore, ablation experiment and replacement experiment further corroborate that the joint understanding of multimodal information can significantly improve the models' performance on proactive embodied action sequence planning task. Our dataset and code are publicly available 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
- Building Cooperative Embodied Agents Modularly with Large Language ModelsHongxin Zhang, Weihua Du, Jiaming Shan, Qinhong Zhou 等ICLR 2024 · 被引用 303 次
- FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM AgentsQinglong Yang, Haoming Li, Haotian Zhao, Xiaokai Yan 等ICLR 2026 · 被引用 21 次
- JointAVBench: A Benchmark for Joint Audio-Visual Reasoning EvaluationJianghan Chao, Jianzhang Gao, Wenhui Tan, Yuchong Sun 等ICLR 2026 · 被引用 16 次
- ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant SimulationJiho Kim, Junseong Choi, Woosog Chay, Daeun Kyung 等ICLR 2026 · 被引用 13 次
相关 Paper
- EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of ThoughtYao Mu, Qinglong Zhang, Mengkang Hu, Wenhai Wang 等NeurIPS 2023 · 被引用 453 次
- Multimodal Embodied Plan Prediction Augmented with Synthetic Embodied DialogueAishwarya Padmakumar, Mert Inan, Spandana Gella, Patrick Lange 等EMNLP 2023 · 被引用 1 次
- Multi-Modal Grounded Planning and Efficient Replanning for Learning Embodied Agents with a Few ExamplesTaewoong Kim, Byeonghwi Kim, Jonghyun ChoiAAAI 2025 · 被引用 8 次
- LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural PlanningShibo Sun, Xue Li, Donglin Di, Mingjie Wei 等ACM MM 2025 · 被引用 4 次
- Steve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open WorldsSipeng Zheng, Jiazheng Liu, Yicheng Feng, Zongqing LuICLR 2024 · 被引用 57 次
