A Mixed Diet Makes DINO An Omnivorous Vision Encoder
Rishabh Kabra, Maks Ovsjanikov, Drew A. Hudson, Ye Xia, Skanda Koppula, André Araújo, João Carreira, Niloy J. Mitra
Abstract
Pre-trained vision encoders like DINOv2 have demonstrated exceptional performance on unimodal tasks. However, we observe that their features are poorly aligned across different visual modalities. For instance, the feature embedding for an RGB image and its corresponding depth map of the same scene exhibit a cosine similarity that is nearly identical to that of two random, unrelated images. To address this, we propose the Omnivorous Vision Encoder, a post-training framework that learns a modality-agnostic feature space. We fine-tune the encoder with a dual objective: first, to maximize the feature alignment between different modalities of the same scene; and second, a distillation objective that anchors the learned representations to a fully frozen teacher. The resulting student encoder becomes "omnivorous" by producing more consistent embeddings for a given scene, regardless of the input modality (RGB, Depth, Segmentation, etc.). This approach enables robust cross-modal understanding while retaining the discriminative semantics of the original foundation model. Omnivorous model weights are available at https://github. com/google-deepmind/representations4d.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0f2559b-8ea1-4264-b50e-39d312866ddeBuilds on22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Perceiver: General Perception with Iterative AttentionAndrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals et al.ICML 2021 · 1,399 citations
- Perceiver IO: A General Architecture for Structured Inputs & OutputsAndrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch et al.ICLR 2022 · 797 citations
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene UnderstandingMike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar et al.ICCV 2021 · 633 citations
- PointOdyssey: A Large-Scale Synthetic Dataset for Long-Term Point TrackingYang Zheng, Adam W. Harley, Bokui Shen, Gordon Wetzstein et al.ICCV 2023 · 255 citations
Related papers
- Omnivore: A Single Model for Many Visual ModalitiesRohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens van der Maaten et al.CVPR 2022 · 185 citations
- Brewing Stronger Features: Dual-Teacher Distillation for Multispectral Earth ObservationFilip Wolf, Blaz Rolih, Luka Cehovin ZajcCVPR 2026 · 4 citations
- OmniVL: One Foundation Model for Image-Language and Video-Language TasksJunke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo et al.NeurIPS 2022 · 205 citations
- Uni-Perceiver: Pre-training Unified Architecture for Generic Perception for Zero-shot and Few-shot TasksXizhou Zhu, Jinguo Zhu, Hao Li, Xiaoshi Wu et al.CVPR 2022
- OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic SegmentationBowen Yin, Jiao-Long Cao, Xuying Zhang, Yuming Chen et al.NeurIPS 2025 · 8 citations
