Learning to Ground Instructional Articles in Videos through Narrations
Effrosyni Mavroudi, Triantafyllos Afouras, Lorenzo Torresani
Abstract
In this paper we present an approach for localizing steps of procedural activities in narrated how-to videos. To deal with the scarcity of labeled data at scale, we source the step descriptions from a language knowledge base (wiki-How) containing instructional articles for a large variety of procedural tasks. Without any form of manual supervision, our model learns to temporally ground the steps of procedural articles in how-to videos by matching three modalities: frames, narrations, and step descriptions. Specifically, our method aligns steps to video by fusing information from two distinct pathways: i) direct alignment of step descriptions to frames, ii) indirect alignment obtained by composing steps-to-narrations with narrations-to-video correspondences. Notably, our approach performs global temporal grounding of all steps in an article at once by exploiting order information, and is trained with step pseudo-labels which are iteratively refined and aggressively filtered. In order to validate our model we introduce a new evaluation benchmark -HT-Step -obtained by manually annotating a 124-hour subset of HowTo100M 1 with steps sourced from wikiHow articles. Experiments on this benchmark as well as zero-shot evaluations on CrossTask demonstrate that our multi-modality alignment yields dramatic gains over several baselines and prior works. Finally, we show that our inner module for matching narration-to-video outperforms by a large margin the state of the art on the HTM-Align narration-video alignment benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3c37f49e-6f5e-46fd-abd3-36a72ea42a1aCited by top-tier papers18
- PerceptionLM: Open-Access Data and Models for Detailed Visual UnderstandingJang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi, Triantafyllos Afouras et al.NeurIPS 2025 · 97 citations
- MatchTime: Towards Automatic Soccer Game Commentary GenerationJiayuan Rao, Haoning Wu, Chang Liu, Yanfeng Wang et al.EMNLP 2024 · 14 citations
- Learning Object State Changes in Videos: An Open-World PerspectiveZihui Xue, Kumar Ashutosh, Kristen GraumanCVPR 2024 · 12 citations
- Step Differences in Instructional VideoTushar Nagarajan, Lorenzo TorresaniCVPR 2024 · 7 citations
- Learning Skill-Attributes for Transferable Assessment in VideoKumar Ashutosh, Kristen GraumanNeurIPS 2025 · 6 citations
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
Related papers
- Learning To Recognize Procedural Activities with Distant SupervisionXudong Lin, Fabio Petroni, Gedas Bertasius, Marcus Rohrbach et al.CVPR 2022 · 55 citations
- Non-Sequential Graph Script Induction via Multimedia GroundingYu Zhou, Sha Li, Manling Li, Xudong Lin et al.ACL 2023 · 5 citations
- Temporal Alignment Networks for Long-term VideoTengda Han, Weidi Xie, Andrew ZissermanCVPR 2022 · 60 citations
- What, When, and Where? Self-Supervised Spatio- Temporal Grounding in Untrimmed Multi-Action Videos from Narrated InstructionsBrian Chen, Nina Shvetsova, Andrew Rouditchenko, Daniel Kondermann et al.CVPR 2024
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
