SVIP: Sequence VerIfication for Procedures in Videos
Yicheng Qian, Weixin Luo, Dongze Lian, Xu Tang, Peilin Zhao, Shenghua Gao
摘要
In this paper, we propose a novel sequence verification task that aims to distinguish positive video pairs performing the same action sequence from negative ones with step-level transformations but still conducting the same task. Such a challenging task resides in an open-set setting without prior action detection or segmentation that requires event-level or even frame-level annotations. To that end, we carefully reorganize two publicly available action-related datasets with step-procedure-task structure. To fully investigate the effectiveness of any method, we collect a scripted video dataset enumerating all kinds of step-level transformations in chemical experiments. Besides, a novel evaluation metric Weighted Distance Ratio is introduced to ensure equivalence for different step-level transformations during evaluation. In the end, a simple but effective baseline based on the transformer encoder with a novel sequence alignment loss is introduced to better characterize long-term dependency between steps, which outperforms other action recognition methods. Codes and data will be released <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> :
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Weakly-Supervised Action Segmentation and Unseen Error Detection in Anomalous Instructional VideosReza Ghoddoosian, Isht Dwivedi, Nakul Agarwal, Behzad DariushICCV 2023 · 被引用 35 次
- OPEL: Optimal Transport Guided ProcedurE LearningSayeed Shafayet Chowdhury, Soumyadeep Chandra, Kaushik RoyNeurIPS 2024 · 被引用 11 次
- Collaborative Weakly Supervised Video Correlation Learning for Procedure-Aware Instructional Video AnalysisTianyao He, Huabin Liu, Yuxi Li, Xiao Ma 等AAAI 2024 · 被引用 8 次
- Error Recognition in Procedural Videos Using Generalized Task GraphShih-Po Lee, Ehsan ElhamifarICCV 2025 · 被引用 3 次
- VideoSetDiff: Identifying and Reasoning Similarities and Differences in Similar VideosYue Qiu, Yanjun Sun, Takuma Yagi, Shusaku Egami 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
相关 Paper
- Weakly Supervised Video Representation Learning with Unaligned Text for Sequential VideosSixun Dong, Huazhang Hu, Dongze Lian, Weixin Luo 等CVPR 2023
- Opening the Vocabulary of Egocentric ActionsDibyadip Chatterjee, Fadime Sener, Shugao Ma, Angela YaoNeurIPS 2023 · 被引用 28 次
- PHGC: Procedural Heterogeneous Graph Completion for Natural Language Task Verification in Egocentric VideosXun Jiang, Zhiyi Huang, Xing Xu, Jingkuan Song 等CVPR 2025
- Learning Procedural-Aware Video Representations Through State-Grounded Hierarchy UnfoldingJinghan Zhao, Yifei Huang, Feng LuAAAI 2026
- On the Importance of Spatial Relations for Few-shot Action RecognitionYilun Zhang, Yuqian Fu, Xingjun Ma, Lizhe Qi 等ACM MM 2023 · 被引用 20 次
