Event-Guided Procedure Planning from Instructional Videos with Text Supervision
An-Lan Wang, Kun-Yu Lin, Jia-Run Du, Jingke Meng, Wei-Shi Zheng
摘要
In this work, we focus on the task of procedure planning from instructional videos with text supervision, where a model aims to predict an action sequence to transform the initial visual state into the goal visual state. A critical challenge of this task is the large semantic gap between observed visual states and unobserved intermediate actions, which is ignored by previous works. Specifically, this semantic gap refers to that the contents in the observed visual states are semantically different from the elements of some action text labels in a procedure. To bridge this semantic gap, we propose a novel event-guided paradigm, which first infers events from the observed states and then plans out actions based on both the states and predicted events. Our inspiration comes from that planning a procedure from an instructional video is to complete a specific event and a specific event usually involves specific actions. Based on the proposed paradigm, we contribute an Event-guided Prompting-based Procedure Planning (E3P) model, which encodes event information into the sequential modeling process to support procedure planning. To further consider the strong action associations within each event, our E3P adopts a mask-and-predict approach for relation mining, incorporating a probabilistic masking scheme for regularization. Extensive experiments on three datasets demonstrate the effectiveness of our proposed model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional VideosYulei Niu, Wenliang Guo, Long Chen, Xudong Lin 等ICLR 2024 · 被引用 26 次
- Single-View Scene Point Cloud Human Grasp GenerationYan-Kang Wang, Chengyi Xing, Yi-Lin Wei, Xiao-Ming Wu 等CVPR 2024 · 被引用 11 次
- PlanLLM: Video Procedure Planning with Refinable Large Language ModelsDejie Yang, Zijing Zhao, Yang LiuAAAI 2025 · 被引用 8 次
- GeoWorld: Geometric World ModelsZeyu Zhang, Danning Li, Ian Reid, Richard HartleyCVPR 2026 · 被引用 6 次
- Why Not Use Your Textbook? Knowledge-Enhanced Procedure Planning of Instructional VideosKumaranage Ravindu Yasas Nagasinghe, Honglu Zhou, Malitha Gunawardhana, Martin Renqiang Min 等CVPR 2024 · 被引用 5 次
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- VidTr: Video Transformer Without ConvolutionsYanyi Zhang, Xinyu Li, Chunhui Liu, Bing Shuai 等ICCV 2021 · 被引用 224 次
相关 Paper
- Masked Temporal Interpolation Diffusion for Procedure Planning in Instructional VideosYufan Zhou, Zhaobo Qi, Lingshuai Lin, Junqi Jing 等ICLR 2025
- P3IV: Probabilistic Procedure Planning from Instructional Videos with Weak SupervisionHe Zhao, Isma Hadji, Nikita Dvornik, Konstantinos G. Derpanis 等CVPR 2022 · 被引用 23 次
- PDPP: Projected Diffusion for Procedure Planning in Instructional VideosHanlin Wang, Yilu Wu, Sheng Guo, Limin WangCVPR 2023
- VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video PromptingMuhammet Furkan Ilaslan, Ali Köksal, Kevin Qinghong Lin, Burak Satar 等AAAI 2025 · 被引用 3 次
- Skip-Plan: Procedure Planning in Instructional Videos via Condensed Action Space LearningZhiheng Li, Wenjia Geng, Muheng Li, Lei Chen 等ICCV 2023 · 被引用 16 次
