Learning Procedural-Aware Video Representations Through State-Grounded Hierarchy Unfolding
Jinghan Zhao, Yifei Huang, Feng Lu
摘要
Learning procedural-aware video representations is a key step towards building agents that can reason about and execute complex tasks. Existing methods typically address this problem by aligning visual content with textual descriptions at the task and step levels to inject procedural semantics into video representations. However, due to their high level of abstraction, "task" and "step" descriptions fail to form a robust alignment with the concrete, observable details in visual data. To address this, we introduce "states", i.e., textual snapshots of object configurations, as a visually-grounded semantic layer that anchors abstract procedures to what a model can actually see. We formalize this insight in a novel Task-Step-State (TSS) framework, where tasks are achieved via steps that drive transitions between observable states. To enforce this structure, we propose a progressive pre-training strategy that unfolds the TSS hierarchy, forcing the model to first ground representations in states before associating them with steps and, ultimately, high-level tasks. Extensive experiments on the COIN and CrossTask datasets show that our method outperforms baseline models on multiple downstream tasks, including task recognition, step recognition, and next step prediction. Ablation studies show that introducing state supervision is a key driver of performance gains across all tasks. Additionally, our progressive pretraining strategy proves more effective than standard joint training, as it better enforces the intended hierarchical structure.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper19
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- Learning To Recognize Procedural Activities with Distant SupervisionXudong Lin, Fabio Petroni, Gedas Bertasius, Marcus Rohrbach 等CVPR 2022 · 被引用 55 次
- Learning to Ground Instructional Articles in Videos through NarrationsEffrosyni Mavroudi, Triantafyllos Afouras, Lorenzo TorresaniICCV 2023 · 被引用 28 次
相关 Paper
- SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional VideosYulei Niu, Wenliang Guo, Long Chen, Xudong Lin 等ICLR 2024 · 被引用 26 次
- Procedure-Aware Pretraining for Instructional Video UnderstandingHonglu Zhou, Roberto Martín-Martín, Mubbasir Kapadia, Silvio Savarese 等CVPR 2023
- VEDIT: Latent Prediction Architecture For Procedural Video Representation LearningHan Lin, Tushar Nagarajan, Nicolas Ballas, Mido Assran 等ICLR 2025
- Skip-Plan: Procedure Planning in Instructional Videos via Condensed Action Space LearningZhiheng Li, Wenjia Geng, Muheng Li, Lei Chen 等ICCV 2023 · 被引用 16 次
- State-aware Video Procedural CaptioningTaichi Nishimura, Atsushi Hashimoto, Yoshitaka Ushiku, Hirotaka Kameko 等ACM MM 2021 · 被引用 15 次
