Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening
Piyush Bagad, Andrew Zisserman
摘要
Our objective is to develop compact video representations that are sensitive to visual change over time. To measure such time-sensitivity, we introduce a new task: chiral action recognition, where one needs to distinguish between a pair of temporally opposite actions, such as "opening vs. closing a door", "approaching vs. moving away from something", "folding vs. unfolding paper", etc. Such actions (i) occur frequently in everyday life, (ii) require understanding of simple visual change over time (in object state, size, spatial position, count . . . ), and (iii) are known to be poorly represented by many video embeddings. Our goal is to build time aware video representations which offer linear separability between these chiral pairs. To that end, we propose a self-supervised adaptation recipe to inject time-sensitivity into a sequence of frozen image features. Our model is based on an auto-encoder with a latent space with inductive bias inspired by perceptual straightening. We show that this results in a compact but time-sensitive video representation for the proposed task across three datasets: Something-Something, EPIC-Kitchens, and Charade. Our method (i) outperforms much larger video models pre-trained on large-scale video datasets, and (ii) leads to an improvement in classification performance on standard benchmarks when combined with these existing models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Seeing the Arrow of Time in Large Multimodal ModelsZihui Xue, Romy Luo, Kristen GraumanNeurIPS 2025 · 被引用 30 次
- Temporal Straightening for Latent PlanningYing Wang, Oumayma Bounou, Gaoyue Zhou, Randall Balestriero 等ICML 2026 · 被引用 19 次
- The Human Brain as a Dynamic Mixture of Expert Models in Video UnderstandingChristina Sartzetaki, Anne Zonneveld, Pablo Oyarzo, Alessandro T. Gifford 等ICLR 2026 · 被引用 4 次
- Frame2Freq: Spectral Adapters for Fine-Grained Video UnderstandingThinesh Thiyakesan Ponbagavathi, Constantin Seibold, Alina RoitbergCVPR 2026 · 被引用 2 次
- From Static to Dynamic: Exploring Self-supervised Image-to-Video Representation Transfer LearningYang Liu, Qianqian Xu, Peisong Wen, Siran Dai 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper53
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
相关 Paper
- Ego-Only: Egocentric Action Detection without Exocentric TransferringHuiyu Wang, Mitesh Kumar Singh, Lorenzo TorresaniICCV 2023 · 被引用 41 次
- Learning State-Aware Visual Representations from Audible InteractionsHimangi Mittal, Pedro Morgado, Unnat Jain, Abhinav GuptaNeurIPS 2022 · 被引用 30 次
- Contrast and Order Representations for Video Self-supervised LearningKai Hu, Jie Shao, Yuan Liu, Bhiksha Raj 等ICCV 2021 · 被引用 76 次
- Ego-Exo: Transferring Visual Representations From Third-Person to First-Person VideosYanghao Li, Tushar Nagarajan, Bo Xiong, Kristen GraumanCVPR 2021
- Learning Spatio-temporal Representation by Channel Aliasing Video PerceptionYiqi Lin, Jinpeng Wang, Manlin Zhang, Andy J. MaACM MM 2021 · 被引用 2 次
