VidLA: Video-Language Alignment at Scale
Mamshad Nayeem Rizve, Fan Fei, Jayakrishnan Unnikrishnan, Son Tran, Benjamin Z. Yao, Belinda Zeng, Mubarak Shah, Trishul Chilimbi
摘要
In this paper, we propose VidLA, an approach for video-language alignment at scale. There are two major limitations of previous video-language alignment approaches. First, they do not capture both short-range and long-range temporal dependencies and typically employ complex hierarchical deep network architectures that are hard to integrate with existing pretrained image-text foundation models. To effectively address this limitation, we instead keep the network architecture simple and use a set of data tokens that operate at different temporal resolutions in a hierarchical manner, accounting for the temporally hierarchical nature of videos. By employing a simple two-tower architecture, we are able to initialize our video-language model with pretrained image-text foundation models, thereby boosting the final performance. Second, existing video-language alignment works struggle due to the lack of semantically aligned large-scale training data. To overcome it, we leverage recent LLMs to curate the largest video-language dataset to date with better visual grounding. Furthermore, unlike existing video-text datasets which only contain short clips, our dataset is enriched with video clips of varying durations to aid our temporally hierarchical data to-kens in extracting better representations at varying temporal scales. Overall, empirical results show that our proposed approach surpasses state-of-the-art methods on Multiple retrieval benchmarks, especially on longer videos, and performs competitively on classification benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Rethinking Video-Language Model from the Language Input PerspectiveXiang Fang, Wanlong Fang, Changshuo Wang, Xiaoye Qu 等AAAI 2026 · 被引用 2 次
- ViLL-E: Video LLM Embeddings for RetrievalRohit Gupta, Jayakrishnan Unnikrishnan, Fan Fei, Sheng Liu 等ACL 2026
- VITRIX-UniViTAR: Unified Vision Transformer with Native ResolutionLimeng Qiao, Yiyang Gan, Bairui Wang, Jie Qin 等NeurIPS 2025
- Video-ColBERT: Contextualized Late Interaction for Text-to-Video RetrievalArun V. Reddy, Alexander Martin, Eugene Yang, Andrew Yates 等CVPR 2025
- LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long VideosTiantian Geng, Jinrui Zhang, Qingni Wang, Teng Wang 等CVPR 2025
它引用的顶会 Paper53
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- HiTeA: Hierarchical Temporal Alignment for Training-Free Long-Video Temporal GroundingXinyi Xu, Hongsong Wang, Guo-Sen Xie, Caifeng Shan 等ICLR 2026
- VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video ModelsHaojian Huang, Haodong Chen, Shengqiong Wu, Meng Luo 等ICML 2025
- Scaling Up Video Summarization Pretraining with Large Language ModelsDawit Mureja Argaw, Seunghyun Yoon, Fabian Caba Heilbron, Hanieh Deilamsalehy 等CVPR 2024
- Learning Beyond Still Frames: Scaling Vision-Language Models with VideoYiyuan Zhang, Handong Jing, Jing Liu, Xiangyu YueICCV 2025 · 被引用 2 次
- Long-Form Video-Language Pre-Training with Multimodal Temporal Contrastive LearningYuchong Sun, Hongwei Xue, Ruihua Song, Bei Liu 等NeurIPS 2022 · 被引用 91 次
