Drop-DTW: Aligning Common Signal Between Sequences While Dropping Outliers
Nikita Dvornik, Isma Hadji, Konstantinos G. Derpanis, Animesh Garg, Allan D. Jepson
摘要
In this work, we consider the problem of sequence-to-sequence alignment for signals containing outliers. Assuming the absence of outliers, the standard Dynamic Time Warping (DTW) algorithm efficiently computes the optimal alignment between two (generally) variable-length sequences. While DTW is robust to temporal shifts and dilations of the signal, it fails to align sequences in a meaningful way in the presence of outliers that can be arbitrarily interspersed in the sequences. To address this problem, we introduce Drop-DTW, a novel algorithm that aligns the common signal between the sequences while automatically dropping the outlier elements from the matching. The entire procedure is implemented as a single dynamic program that is efficient and fully differentiable. In our experiments, we show that Drop-DTW is a robust similarity measure for sequence retrieval and demonstrate its effectiveness as a training loss on diverse applications. With Drop-DTW, we address temporal step localization on instructional videos, representation learning from noisy videos, and cross-modal representation learning for audio-visual retrieval and localization. In all applications, we take a weakly- or unsupervised approach and demonstrate state-of-the-art results under these settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Fine-grained Temporal Contrastive Learning for Weakly-supervised Temporal Action LocalizationJunyu Gao, Mengyuan Chen, Changsheng XuCVPR 2022 · 被引用 87 次
- Temporal Alignment Networks for Long-term VideoTengda Han, Weidi Xie, Andrew ZissermanCVPR 2022 · 被引用 60 次
- Video-Mined Task Graphs for Keystep Recognition in Instructional VideosKumar Ashutosh, Santhosh Kumar Ramakrishnan, Triantafyllos Afouras, Kristen GraumanNeurIPS 2023 · 被引用 51 次
- Multi-granularity Correspondence Learning from Long-term Noisy VideosYijie Lin, Jie Zhang, Zhenyu Huang, Jia Liu 等ICLR 2024 · 被引用 42 次
- Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric VideosLuigi Seminara, Giovanni Maria Farinella, Antonino FurnariNeurIPS 2024 · 被引用 36 次
它引用的顶会 Paper9
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- Space-Time Correspondence as a Contrastive Random WalkAllan Jabri, Andrew Owens, Alexei A. EfrosNeurIPS 2020 · 被引用 356 次
- Weakly-Supervised Action Localization With Background ModelingPhuc Xuan Nguyen, Deva Ramanan, Charless C. FowlkesICCV 2019 · 被引用 176 次
- Representation Learning via Global Temporal Alignment and Cycle-ConsistencyIsma Hadji, Konstantinos G. Derpanis, Allan D. JepsonCVPR 2021
相关 Paper
- Learning To Segment Actions From Visual and Language Instructions via Differentiable Weak Sequence AlignmentYuhan Shen, Lu Wang, Ehsan ElhamifarCVPR 2021
- TempCLR: Temporal Alignment Representation with Contrastive LearningYuncong Yang, Jiawei Ma, Shiyuan Huang, Long Chen 等ICLR 2023
- Video-Text Representation Learning via Differentiable Weak Temporal AlignmentDohwan Ko, Joonmyung Choi, Juyeon Ko, Shinyeong Noh 等CVPR 2022 · 被引用 18 次
- Exploring Denoised Cross-video Contrast for Weakly-supervised Temporal Action LocalizationJingjing Li, Tianyu Yang, Wei Ji, Jue Wang 等CVPR 2022 · 被引用 57 次
- Weakly Supervised Temporal Anomaly Segmentation with Dynamic Time WarpingDongha Lee, Sehun Yu, Hyunjun Ju, Hwanjo YuICCV 2021 · 被引用 17 次
