AM-RADIO: Agglomerative Vision Foundation Model Reduce All Domains Into One
Mike Ranzinger, Greg Heinrich, Jan Kautz, Pavlo Molchanov
Abstract
A handful of visual foundation models (VFMs) have recently emerged as the backbones for numerous downstream tasks. VFMs like CLIP, DINOv2, SAM are trained with distinct objectives, exhibiting unique characteristics for various downstream tasks. We find that despite their conceptual differences, these models can be effectively merged into a unified model through multi-teacher distillation. We name this approach AM-RADIO (Agglomerative Model - Reduce All Domains Into One). This integrative approach not only surpasses the performance of individual teacher models but also amalgamates their distinctive features, such as zero-shot vision-language comprehension, detailed pixel-level understanding, and open vocabulary segmentation capabilities. Additionally, in pursuit of the most hardware-efficient backbone, we evaluated numerous architectures in our multi-teacher distillation pipeline using the same training recipe. This led to the development of a novel architecture (E-RADIO) that exceeds the performance of its predecessors and is at least 6x faster than the teacher models at matched resolution. Our comprehensive benchmarking process covers downstream tasks including ImageNet classification, semantic segmentation linear probing, COCO object detection and integration into LLaVa-1.5. Code: https://github.com/NVlabs/RADIO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28ee7725-87c8-426a-a8c0-c7f43c80c864Cited by top-tier papers57
- Perception Encoder: The best visual embeddings are not at the output of the networkDaniel Bolya, Po-Yao Huang, Peize Sun, Jang Hyun Cho et al.NeurIPS 2025 · 359 citations
- Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial RepresentationsYujia Zhang, Xiaoyang Wu, Yixing Lao, Chengyao Wang et al.NeurIPS 2025 · 47 citations
- VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action ModelsChongkai Gao, Zixuan Liu, Zhenghao Chi, Junshan Huang et al.NeurIPS 2025 · 41 citations
- SenseFlow: Scaling Distribution Matching for Flow-based Text-to-Image DistillationXingtong Ge, Xin Zhang, Tongda Xu, Yi Zhang et al.ICLR 2026 · 29 citations
- AnyUp: Universal Feature UpsamplingThomas Wimmer, Prune Truong, Marie-Julie Rakotosaona, Michael Oechsle et al.ICLR 2026 · 29 citations
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 4,239 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
Related papers
- RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation ModelsGreg Heinrich, Mike Ranzinger, Hongxu Yin, Yao Lu et al.CVPR 2025
- RADIO1D: Elastic Representations for Condensed Vision ModelingGreg Heinrich, Mike Ranzinger, Collin McCarthy, Natan Bagrov et al.ICML 2026
- SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation ModelsSofian Chaybouti, Sanath Narayan, Yasser Dahou, Phúc H. Lê Khắc et al.CVPR 2026 · 2 citations
- Building Vision-Language Models on Solid Foundations with Masked DistillationSepehr Sameni, Kushal Kafle, Hao Tan, Simon JenniCVPR 2024 · 4 citations
- TransAgent: Transfer Vision-Language Foundation Models with Heterogeneous Agent CollaborationYiwei Guo, Shaobin Zhuang, Kunchang Li, Yu Qiao et al.NeurIPS 2024 · 9 citations
