DETACH : Decomposed Spatio-Temporal Alignment for Exocentric Video and Ambient Sensors with Staged Learning
Junho Yoon, Jaemo Jeong, Hyunju Kim, Dongman Lee
摘要
Aligning egocentric video with wearable sensors have shown promise for human action recognition, but face practical limitations in user discomfort, privacy concerns, and scalability. We explore exocentric video with ambient sensors as a non-intrusive, scalable alternative. While prior egocentric-wearable works predominantly adopt Global Alignment by encoding entire sequences into unified representations, this approach fails in exocentric-ambient settings due to two problems: (P1) inability to capture local details such as subtle motions, and (P2) over-reliance on modality-invariant temporal patterns, causing misalignment between actions sharing similar temporal patterns with different spatio-semantic contexts. To resolve these problems, we propose DETACH, a decomposed spatio-temporal framework. This explicit decomposition preserves local details, while our novel sensor-spatial features discovered via online clustering provide semantic grounding for context-aware alignment. To align the decomposed features, our two-stage approach establishes spatial correspondence through mutual supervision, then performs temporal alignment via a spatial-temporal weighted contrastive loss that adaptively handles easy negatives, hard negatives, and false negatives. Comprehensive experiments with downstream tasks on Opportunity++ and HWU-USP datasets demonstrate substantial improvements over adapted egocentric-wearable baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
相关 Paper
- MoBind: Motion Binding for Fine-Grained IMU-Video Pose AlignmentDuc Duy Nguyen, Tat-Jun Chin, Minh HoaiCVPR 2026 · 被引用 1 次
- Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal AlignmentZihui Xue, Kristen GraumanNeurIPS 2023 · 被引用 64 次
- Ego-Only: Egocentric Action Detection without Exocentric TransferringHuiyu Wang, Mitesh Kumar Singh, Lorenzo TorresaniICCV 2023 · 被引用 41 次
- COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity RecognitionBaiyu Chen, Wilson Wongso, Zechen Li, Yonchanok Khaokaew 等UbiComp 2026 · 被引用 1 次
- WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity RecognitionMarius Bock, Hilde Kuehne, Kristof Van Laerhoven, Michael MöllerUbiComp 2025 · 被引用 50 次
