Dynamic Reflections: Probing Video Representations with Text Alignment
Maks Ovsjanikov, Viorica Patraucean, Leonidas J. Guibas, Tyler Zhu, Tengda Han
摘要
The alignment of representations from different modalities has recently been shown to provide insights on the structural similarities and downstream capabilities of different encoders across diverse data types. While significant progress has been made in aligning images with text, the temporal nature of video data remains largely unexplored in this context. In this work, we conduct the first comprehensive study of video-text representation alignment, probing the capabilities of modern video and language encoders. Our findings reveal several key insights. First, we demonstrate that cross-modal alignment highly depends on the richness of both visual (static images vs. multi-frame videos) and text (single caption vs. a collection) data provided at test time, especially when using state-of-the-art video encoders. We propose parametric test-time scaling laws that capture this behavior and show remarkable predictive power against empirical observations. Secondly, we investigate the correlation between semantic alignment and performance on both semantic and non-semantic downstream tasks, providing initial evidence that strong alignment against text encoders may be linked to general-purpose video representation and understanding. Finally, we correlate temporal reasoning with cross-modal alignment providing a challenging test-bed for vision and language models. Overall, our work introduces video-text alignment as an informative zero-shot way to probe the representation power of different encoders for spatio-temporal data. Project page can be found at https://video-prh.github.io/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Revisiting the Platonic Representation Hypothesis: An Aristotelian ViewFabian Gröger, Shuo Wen, Maria BrbicICML 2026 · 被引用 27 次
- The Truth Lies Somewhere in the Middle (of the Generated Tokens)Sophie Wang, Phillip Isola, Brian CheungICML 2026 · 被引用 1 次
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li 等ICCV 2019 · 被引用 688 次
相关 Paper
- Time Series, Vision, and Language: Exploring the Limits of Alignment in Contrastive Representation SpacesPratham Yashwante, Rose YuICML 2026
- VidLA: Video-Language Alignment at ScaleMamshad Nayeem Rizve, Fan Fei, Jayakrishnan Unnikrishnan, Son Tran 等CVPR 2024 · 被引用 3 次
- HiTeA: Hierarchical Temporal-Aware Video-Language Pre-trainingQinghao Ye, Guohai Xu, Ming Yan, Haiyang Xu 等ICCV 2023 · 被引用 102 次
- Align and Prompt: Video-and-Language Pre-training with Entity PromptsDongxu Li, Junnan Li, Hongdong Li, Juan Carlos Niebles 等CVPR 2022
- Temporal Reasoning Transfer from Text to VideoLei Li, Yuanxin Liu, Linli Yao, Peiyuan Zhang 等ICLR 2025
