Procedure Planning in Instructional Videos via Contextual Modeling and Model-based Policy Learning
Jing Bi, Jiebo Luo, Chenliang Xu
摘要
Learning new skills by observing humans’ behaviors is an essential capability of AI. In this work, we leverage instructional videos to study humans’ decision-making processes, focusing on learning a model to plan goal-directed actions in real-life videos. In contrast to conventional action recognition, goal-directed actions are based on expectations of their outcomes requiring causal knowledge of potential consequences of actions. Thus, integrating the environment structure with goals is critical for solving this task. Previous works learn a single world model will fail to distinguish various tasks, resulting in an ambiguous la-tent space; planning through it will gradually neglect the desired outcomes since the global information of the future goal degrades quickly as the procedure evolves. We address these limitations with a new formulation of procedure planning and propose novel algorithms to model human behaviors through Bayesian Inference and model-based Imitation Learning. Experiments conducted on real-world instructional videos show that our method can achieve state-of-the-art performance in reaching the indicated goals. Furthermore, the learned contextual information presents interesting features for planning in a latent space.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?Qi Zhao, Shijie Wang, Ce Zhang, Changcheng Fu 等ICLR 2024 · 被引用 93 次
- Video-Mined Task Graphs for Keystep Recognition in Instructional VideosKumar Ashutosh, Santhosh Kumar Ramakrishnan, Triantafyllos Afouras, Kristen GraumanNeurIPS 2023 · 被引用 51 次
- Pretrained Language Models as Visual Planners for Human AssistanceDhruvesh Patel, Hamid Eghbalzadeh, Nitin Kamra, Michael Louis Iuzzolino 等ICCV 2023 · 被引用 41 次
- Learning to Ground Instructional Articles in Videos through NarrationsEffrosyni Mavroudi, Triantafyllos Afouras, Lorenzo TorresaniICCV 2023 · 被引用 28 次
- SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional VideosYulei Niu, Wenliang Guo, Long Chen, Xudong Lin 等ICLR 2024 · 被引用 26 次
它引用的顶会 Paper3
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski 等ICLR 2020 · 被引用 969 次
- Learning to Paint With Model-Based Deep Reinforcement LearningZhewei Huang, Shuchang Zhou, Wen HengICCV 2019 · 被引用 180 次
- Action Recognition With Spatial-Temporal Discriminative Filter BanksBrais Martínez, Davide Modolo, Yuanjun Xiong, Joseph TigheICCV 2019 · 被引用 70 次
相关 Paper
- Learning Latent Action World Models in the WildQuentin Garrido, Tushar Nagarajan, Basile Terver, Nicolas Ballas 等ICML 2026 · 被引用 38 次
- Planning from Pixels using Inverse Dynamics ModelsKeiran Paster, Sheila A. McIlraith, Jimmy BaICLR 2021 · 被引用 44 次
- P3IV: Probabilistic Procedure Planning from Instructional Videos with Weak SupervisionHe Zhao, Isma Hadji, Nikita Dvornik, Konstantinos G. Derpanis 等CVPR 2022 · 被引用 23 次
- Skip-Plan: Procedure Planning in Instructional Videos via Condensed Action Space LearningZhiheng Li, Wenjia Geng, Muheng Li, Lei Chen 等ICCV 2023 · 被引用 16 次
- AdaWorld: Learning Adaptable World Models with Latent ActionsShenyuan Gao, Siyuan Zhou, Yilun Du, Jun Zhang 等ICML 2025
