Detours for Navigating Instructional Videos
Kumar Ashutosh, Zihui Xue, Tushar Nagarajan, Kristen Grauman
摘要
We introduce the video detours problem for navigating instructional videos. Given a source video and a natural language query asking to alter the how-to video's current path of execution in a certain way, the goal is to find a related "detour video" that satisfies the requested alteration. To address this challenge, we propose VidDetours, a novel video-language approach that learns to retrieve the targeted temporal segments from a large repository of how-to's using video-and-text conditioned queries. Furthermore, we devise a language-based pipeline that exploits how-to video narration text to create weakly supervised training data. We demonstrate our idea applied to the domain of how-to cooking videos, where a user can detour from their current recipe to find steps with alternate ingredients, tools, and techniques. Validating on a ground truth annotated dataset of 16K samples, we show our model's significant improvements over best available methods for video retrieval and question answering, with recall rates exceeding the state of the art by 35%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Vid2Coach: Transforming How-To Videos into Task AssistantsMina Huh, Zihui Xue, Ujjaini Das, Kumar Ashutosh 等UIST 2025 · 被引用 9 次
- Step Differences in Instructional VideoTushar Nagarajan, Lorenzo TorresaniCVPR 2024 · 被引用 7 次
- Learning Skill-Attributes for Transferable Assessment in VideoKumar Ashutosh, Kristen GraumanNeurIPS 2025 · 被引用 6 次
- Stitch-a-Demo: Creating Video Demonstrations from Multistep DescriptionsChi Hsuan Wu, Kumar Ashutosh, Kristen GraumanCVPR 2026 · 被引用 1 次
- ExpertAF: Expert Actionable Feedback from VideoKumar Ashutosh, Tushar Nagarajan, Georgios Pavlakos, Kris Kitani 等CVPR 2025
它引用的顶会 Paper44
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
相关 Paper
- Composed Video Retrieval via Enriched Context and Discriminative EmbeddingsOmkar Thawakar, Muzammal Naseer, Rao Muhammad Anwer, Salman H. Khan 等CVPR 2024 · 被引用 10 次
- Temporal Alignment Networks for Long-term VideoTengda Han, Weidi Xie, Andrew ZissermanCVPR 2022 · 被引用 60 次
- ProTéGé: Untrimmed Pretraining for Video Temporal Grounding by Video Temporal GroundingLan Wang, Gaurav Mittal, Sandra Sajeev, Ye Yu 等CVPR 2023
- Look at What I'm Doing: Self-Supervised Spatial Grounding of Narrations in Instructional VideosReuben Tan, Bryan A. Plummer, Kate Saenko, Hailin Jin 等NeurIPS 2021 · 被引用 30 次
- Video-Text Representation Learning via Differentiable Weak Temporal AlignmentDohwan Ko, Joonmyung Choi, Juyeon Ko, Shinyeong Noh 等CVPR 2022 · 被引用 18 次
