VicTR: Video-conditioned Text Representations for Activity Recognition
Kumara Kahatapitiya, Anurag Arnab, Arsha Nagrani, Michael S. Ryoo
Abstract
Vision-Language models (VLMs) have excelled in the image-domain-especially in zero-shot settings-thanks to the availability of vast pretraining data (i.e., paired imagetext samples). However for videos, such paired data is not as abundant. Therefore, video-VLMs are usually designed by adapting pretrained image-VLMs to the video-domain, instead of training from scratch. All such recipes rely on augmenting visual embeddings with temporal information (i.e., image → video), often keeping text embeddings unchanged or even being discarded. In this paper, we argue the contrary, that better video-VLMs can be designed by focusing more on augmenting text, rather than visual information. More specifically, we introduce Video-conditioned Text Representations (VicTR): a form of text embeddings optimized w.r.t. visual embeddings, creating a more-flexible contrastive latent space. Our model can further make use of freely-available semantic information, in the form of visually-grounded auxiliary text (e.g. object or scene information). We evaluate our model on few-shot, zero-shot (HMDB-51, UCF-101), short-form (Kinetics-400) and long-form (Charades) activity recognition benchmarks, showing strong performance among video-VLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ebfc5e22-a76f-431c-92e2-4ad431a3cac9Cited by top-tier papers10
- Time Blindness: Why Video-Language Models Can't See What Humans Can?Ujjwal Upadhyay, Mukul Ranjan, Zhiqiang Shen, Mohamed ElhoseinyCVPR 2026 · 17 citations
- Language-based Action Concept Spaces Improve Video Self-Supervised LearningKanchana Ranasinghe, Michael S. RyooNeurIPS 2023 · 16 citations
- Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIPYating Yu, Congqi Cao, Yueran Zhang, Qinyi Lv et al.AAAI 2025 · 12 citations
- Generating Action-conditioned Prompts for Open-vocabulary Video Action RecognitionChengyou Jia, Minnan Luo, Xiaojun Chang, Zhuohang Dang et al.ACM MM 2024 · 10 citations
- VTD-CLIP: Video-to-Text Discretization via Prompting CLIPWencheng Zhu, Yuexin Wang, Hongxuan Li, Pengfei ZhuAAAI 2026 · 2 citations
Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
Related papers
- HierVL: Learning Hierarchical Video-Language EmbeddingsKumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, Kristen GraumanCVPR 2023
- MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language KnowledgeWei Lin, Leonid Karlinsky, Nina Shvetsova, Horst Possegger et al.ICCV 2023 · 52 citations
- OST: Refining Text Knowledge with Optimal Spatio-Temporal Descriptor for General Video RecognitionTom Tongjia Chen, Hongshan Yu, Zhengeng Yang, Zechuan Li et al.CVPR 2024
- Bidirectional Cross-Modal Knowledge Exploration for Video Recognition with Pre-trained Vision-Language ModelsWenhao Wu, Xiaohan Wang, Haipeng Luo, Jingdong Wang et al.CVPR 2023
- Towards Generalisable Video Moment Retrieval: Visual-Dynamic Injection to Image-Text Pre-TrainingDezhao Luo, Jiabo Huang, Shaogang Gong, Hailin Jin et al.CVPR 2023
