Procedure Planning in Instructional Videos via Contextual Modeling and Model-based Policy Learning
Jing Bi, Jiebo Luo, Chenliang Xu
Abstract
Learning new skills by observing humans’ behaviors is an essential capability of AI. In this work, we leverage instructional videos to study humans’ decision-making processes, focusing on learning a model to plan goal-directed actions in real-life videos. In contrast to conventional action recognition, goal-directed actions are based on expectations of their outcomes requiring causal knowledge of potential consequences of actions. Thus, integrating the environment structure with goals is critical for solving this task. Previous works learn a single world model will fail to distinguish various tasks, resulting in an ambiguous la-tent space; planning through it will gradually neglect the desired outcomes since the global information of the future goal degrades quickly as the procedure evolves. We address these limitations with a new formulation of procedure planning and propose novel algorithms to model human behaviors through Bayesian Inference and model-based Imitation Learning. Experiments conducted on real-world instructional videos show that our method can achieve state-of-the-art performance in reaching the indicated goals. Furthermore, the learned contextual information presents interesting features for planning in a latent space.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 25e7d7b1-720b-4be5-8313-795740a9d945Cited by top-tier papers32
- AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?Qi Zhao, Shijie Wang, Ce Zhang, Changcheng Fu et al.ICLR 2024 · 93 citations
- Video-Mined Task Graphs for Keystep Recognition in Instructional VideosKumar Ashutosh, Santhosh Kumar Ramakrishnan, Triantafyllos Afouras, Kristen GraumanNeurIPS 2023 · 51 citations
- Pretrained Language Models as Visual Planners for Human AssistanceDhruvesh Patel, Hamid Eghbalzadeh, Nitin Kamra, Michael Louis Iuzzolino et al.ICCV 2023 · 41 citations
- Learning to Ground Instructional Articles in Videos through NarrationsEffrosyni Mavroudi, Triantafyllos Afouras, Lorenzo TorresaniICCV 2023 · 28 citations
- SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional VideosYulei Niu, Wenliang Guo, Long Chen, Xudong Lin et al.ICLR 2024 · 26 citations
Builds on3
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
- Learning to Paint With Model-Based Deep Reinforcement LearningZhewei Huang, Shuchang Zhou, Wen HengICCV 2019 · 180 citations
- Action Recognition With Spatial-Temporal Discriminative Filter BanksBrais Martínez, Davide Modolo, Yuanjun Xiong, Joseph TigheICCV 2019 · 70 citations
Related papers
- Learning Latent Action World Models in the WildQuentin Garrido, Tushar Nagarajan, Basile Terver, Nicolas Ballas et al.ICML 2026 · 38 citations
- Planning from Pixels using Inverse Dynamics ModelsKeiran Paster, Sheila A. McIlraith, Jimmy BaICLR 2021 · 44 citations
- P3IV: Probabilistic Procedure Planning from Instructional Videos with Weak SupervisionHe Zhao, Isma Hadji, Nikita Dvornik, Konstantinos G. Derpanis et al.CVPR 2022 · 23 citations
- Skip-Plan: Procedure Planning in Instructional Videos via Condensed Action Space LearningZhiheng Li, Wenjia Geng, Muheng Li, Lei Chen et al.ICCV 2023 · 16 citations
- AdaWorld: Learning Adaptable World Models with Latent ActionsShenyuan Gao, Siyuan Zhou, Yilun Du, Jun Zhang et al.ICML 2025
