SViTT: Temporal Learning of Sparse Video-Text Transformers
Yi Li, Kyle Min, Subarna Tripathi, Nuno Vasconcelos
Abstract
Do video-text transformers learn to model temporal relationships across frames? Despite their immense capacity and the abundance of multimodal training data, recent work has revealed the strong tendency of video-text models towards frame-based spatial representations, while temporal reasoning remains largely unsolved. In this work, we identify several key challenges in temporal learning of videotext transformers: the spatiotemporal trade-off from limited network size; the curse of dimensionality for multi-frame modeling; and the diminishing returns of semantic information by extending clip length. Guided by these findings, we propose SViTT, a sparse video-text architecture that performs multi-frame reasoning with significantly lower cost than naïve transformers with dense attention. Analogous to graph-based networks, SViTT employs two forms of sparsity: edge sparsity that limits the query-key communications between tokens in self-attention, and node sparsity that discards uninformative visual tokens. Trained with a curriculum which increases model sparsity with the clip length, SViTT outperforms dense transformer baselines on multiple video-text retrieval and question answering benchmarks, with a fraction of computational cost. Project
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Text Is MASS: Modeling as Stochastic Embedding for Text-Video RetrievalJiamian Wang, Pichao Wang, Guohao Sun, Dongfang Liu et al.CVPR 2024 · 52 citations
- CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge TransferYabing Wang, Fan Wang, Jianfeng Dong, Hao LuoAAAI 2024 · 20 citations
- Diffusion-Inspired Truncated Sampler for Text-Video RetrievalJiamian Wang, Pichao Wang, Dongfang Liu, Qiang Guan et al.NeurIPS 2024 · 16 citations
- One Trajectory, One Token: Grounded Video Tokenization Via Panoptic Sub-Object TrajectoryChenhao Zheng, Jieyu Zhang, Mohammadreza Salehi, Ziqi Gao et al.ICCV 2025 · 7 citations
- Ranking Distillation for Open-Ended Video Question Answering with Insufficient LabelsTianming Liang, Chaolei Tan, Beihao Xia, Wei-Shi Zheng et al.CVPR 2024
Builds on33
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
Related papers
- ResidualViT for Efficient Temporally Dense Video EncodingMattia Soldan, Fabian Caba Heilbron, Bernard Ghanem, Josef Sivic et al.ICCV 2025
- SwinBERT: End-to-End Transformers with Sparse Attention for Video CaptioningKevin Lin, Linjie Li, Chung-Ching Lin, Faisal Ahmed et al.CVPR 2022 · 263 citations
- Less Is More: ClipBERT for Video-and-Language Learning via Sparse SamplingJie Lei, Linjie Li, Luowei Zhou, Zhe Gan et al.CVPR 2021
- VecAttention: Vector-wise Sparse Attention for Accelerating Long Context InferenceAnmin Liu, Ruixuan Yang, Huiqiang Jiang, Bin Lin et al.CVPR 2026 · 4 citations
- M-LLM Based Video Frame Selection for Efficient Video UnderstandingKai Hu, Feng Gao, Xiaohan Nie, Peng Zhou et al.CVPR 2025
