UniViT: Unifying Image and Video Understanding in One Vision Encoder
Feilong Tang, Xiang An, Haolin Yang, Yin Xie, Kaicheng Yang, Ming Hu, Zheng Cheng, Xingyu Zhou, Zimin Ran, Imran Razzak, Ziyong Feng, Behzad Bozorgtabar
Abstract
Despite the impressive progress of recent pretraining methods on multimodal tasks, existing methods are inherently biased towards either spatial modeling ( e.g. , CLIP) or temporal modeling ( e.g. , V-JEPA), limiting their joint capture of spatial details and temporal dynamics. To this end, we propose UniViT , a cluster-driven unified self-supervised learning framework that effectively captures the structured semantics of both image spatial content and video temporal dynamics through event-level and object-level clustering and discrimination. Specifically, we leverage offline clustering to generate semantic clusters across both modalities. For videos, multi-granularity event-level clustering progressively expands from single-event to structured multi-event segments, capturing coarse-to-fine temporal semantics; for images, object-level clustering captures fine-grained spatial semantics. However, while global clustering provides semantically consistent clusters, it lacks modeling of structured semantic relations ( e.g., temporal event structures). To address this, we introduce a contrastive objective that leverages these semantic clusters as pseudo-label supervision to explicitly enforce structural constraints, including temporal event relations and spatial object co-occurrences, capturing structured semantics beyond categories. Meanwhile, UniViT jointly embeds structured object-level and event-level semantics into a unified representation space. Furthermore, UniViT introduces two key components: (i) Unified Rotary Position Embedding integrates relative positional embedding with frequency-aware dimension allocation to support position-invariant semantic learning and enhance the stability of structured semantics in the discrimination stage; and (ii) Variable Spatiotemporal Streams adapt to inputs of varying frame lengths, addressing the rigidity of conventional fixed-input approaches. Extensive experiments across varying model scales demonstrate that UniViT achieves state-of-the-art performance on linear probing, attentive probing, question answering, and spatial understanding tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model PerspectiveHangjie Yuan, Weihua Chen, Jun Cen, Hu Yu et al.ICLR 2026 · 21 citations
- Leveraging Class Distributions in CLIP for Weakly Supervised Semantic SegmentationZiqian Yang, Xinqiao Zhao, Xiaolei Wang, Quan Zhang et al.CVPR 2026
- Frequency-Aware Affinity for Weakly Supervised Semantic SegmentationZiqian Yang, Xianglin Qiu, Xinqiao Zhao, Xiaolei Wang et al.CVPR 2026
- Remember Me: Bridging the Long-Range Gap in LVLMs with Three-Step Inference-Only Decay Resilience StrategiesPeng Gao, Yujian Lee, Xiaofeng Zhang, Zailong Chen et al.AAAI 2026
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
Related papers
- Vi2CLR: Video and Image for Visual Contrastive Learning of RepresentationAli Diba, Vivek Sharma, Reza Safdari, Dariush Lotfi et al.ICCV 2021 · 65 citations
- Point Contrastive Prediction with Semantic Clustering for Self-Supervised Learning on Point Cloud VideosXiaoxiao Sheng, Zhiqiang Shen, Gang Xiao, Longguang Wang et al.ICCV 2023 · 20 citations
- Contextualized Spatio-Temporal Contrastive Learning with Self-SupervisionLiangzhe Yuan, Rui Qian, Yin Cui, Boqing Gong et al.CVPR 2022 · 24 citations
- HiVLP: Hierarchical Interactive Video-Language Pre-TrainingBin Shao, Jianzhuang Liu, Renjing Pei, Songcen Xu et al.ICCV 2023 · 6 citations
- SMILE: Infusing Spatial and Motion Semantics in Masked Video LearningFida Mohammad Thoker, Letian Jiang, Chen Zhao, Bernard GhanemCVPR 2025
