VicTR: Video-conditioned Text Representations for Activity Recognition
Kumara Kahatapitiya, Anurag Arnab, Arsha Nagrani, Michael S. Ryoo
摘要
Vision-Language models (VLMs) have excelled in the image-domain-especially in zero-shot settings-thanks to the availability of vast pretraining data (i.e., paired imagetext samples). However for videos, such paired data is not as abundant. Therefore, video-VLMs are usually designed by adapting pretrained image-VLMs to the video-domain, instead of training from scratch. All such recipes rely on augmenting visual embeddings with temporal information (i.e., image → video), often keeping text embeddings unchanged or even being discarded. In this paper, we argue the contrary, that better video-VLMs can be designed by focusing more on augmenting text, rather than visual information. More specifically, we introduce Video-conditioned Text Representations (VicTR): a form of text embeddings optimized w.r.t. visual embeddings, creating a more-flexible contrastive latent space. Our model can further make use of freely-available semantic information, in the form of visually-grounded auxiliary text (e.g. object or scene information). We evaluate our model on few-shot, zero-shot (HMDB-51, UCF-101), short-form (Kinetics-400) and long-form (Charades) activity recognition benchmarks, showing strong performance among video-VLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Time Blindness: Why Video-Language Models Can't See What Humans Can?Ujjwal Upadhyay, Mukul Ranjan, Zhiqiang Shen, Mohamed ElhoseinyCVPR 2026 · 被引用 17 次
- Language-based Action Concept Spaces Improve Video Self-Supervised LearningKanchana Ranasinghe, Michael S. RyooNeurIPS 2023 · 被引用 16 次
- Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIPYating Yu, Congqi Cao, Yueran Zhang, Qinyi Lv 等AAAI 2025 · 被引用 12 次
- Generating Action-conditioned Prompts for Open-vocabulary Video Action RecognitionChengyou Jia, Minnan Luo, Xiaojun Chang, Zhuohang Dang 等ACM MM 2024 · 被引用 10 次
- VTD-CLIP: Video-to-Text Discretization via Prompting CLIPWencheng Zhu, Yuexin Wang, Hongxuan Li, Pengfei ZhuAAAI 2026 · 被引用 2 次
它引用的顶会 Paper37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
相关 Paper
- HierVL: Learning Hierarchical Video-Language EmbeddingsKumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, Kristen GraumanCVPR 2023
- MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language KnowledgeWei Lin, Leonid Karlinsky, Nina Shvetsova, Horst Possegger 等ICCV 2023 · 被引用 52 次
- OST: Refining Text Knowledge with Optimal Spatio-Temporal Descriptor for General Video RecognitionTom Tongjia Chen, Hongshan Yu, Zhengeng Yang, Zechuan Li 等CVPR 2024
- Bidirectional Cross-Modal Knowledge Exploration for Video Recognition with Pre-trained Vision-Language ModelsWenhao Wu, Xiaohan Wang, Haipeng Luo, Jingdong Wang 等CVPR 2023
- Towards Generalisable Video Moment Retrieval: Visual-Dynamic Injection to Image-Text Pre-TrainingDezhao Luo, Jiabo Huang, Shaogang Gong, Hailin Jin 等CVPR 2023
