Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
Shashanka Venkataramanan, Valentinos Pariza, Mohammadreza Salehi, Lukas Knobel, Elias Ramzi, Spyros Gidaris, Andrei Bursuc, Yuki M Asano
Abstract
We present Franca (pronounced Fran-ka): free one; the first fully open-source (data, code, weights) vision foundation model that matches and in many cases surpasses the performance of state-of-the-art proprietary models, e.g., DINOv2, CLIP, SigLIPv2, etc. Our approach is grounded in a transparent training pipeline inspired by Web-SSL and uses publicly available data: ImageNet-21K and a subset of ReLAION-2B. Beyond model release, we tackle critical limitations in self-supervised learning clustering methods. Existing approaches assign image features to large codebooks via clustering algorithms such as Sinkhorn-Knopp, but they often overlook the inherent ambiguity in cluster semantics. To address this, we introduce a multi-head clustering projector based on nested Matryoshka representations. This design progressively refines features into increasingly fine-grained clusters without increasing the model size, producing higher-quality dense representations. Additionally, we propose a novel positional disentanglement strategy that explicitly removes positional biases from dense representations.This leads to consistent gains on several downstream benchmarks, demonstrating the utility of cleaner feature spaces. Our contributions establish a new standard for transparent, high-performance vision models and open a path toward more reproducible and generalizable foundation models for the broader AI community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7efd0e93-dbca-4ae2-84a0-e55dbe1639faCited by top-tier papers16
- DINO-Foresight: Looking into the Future with DINOEfstathios Karypidis, Ioannis Kakogeorgiou, Spyridon Gidaris, Nikos KomodakisNeurIPS 2025 · 52 citations
- TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text AlignmentBingyi Cao, Koert Chen, Kevis-Kokitsi Maninis, Kaifeng Chen et al.CVPR 2026 · 14 citations
- INSID3: Training-Free In-Context Segmentation with DINOv3Claudia Cuttano, Gabriele Trivigno, Christoph Reich, Daniel Cremers et al.CVPR 2026 · 13 citations
- Attention, Please! Revisiting Attentive Probing Through the Lens of EfficiencyBill Psomas, Dionysis Christopoulos, Eirini Baltzi, Ioannis Kakogeorgiou et al.ICLR 2026 · 12 citations
- LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-SupervisionAnthony Fuller, Yousef Yassin, Junfeng Wen, Tarek Ibrahim et al.NeurIPS 2025 · 7 citations
Builds on44
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
Related papers
- Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT TransformersCris Claessens, Christiaan Viviers, Giacomo D'Amicantonio, Egor Bondarev et al.CVPR 2026 · 6 citations
- Accessing Vision Foundation Models via ImageNet-1KYitian Zhang, Xu Ma, Yue Bai, Huan Wang et al.ICLR 2025
- Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific ModelsRaviteja Vemulapalli, Hadi Pouransari, Fartash Faghri, Sachin Mehta et al.ICML 2024 · 15 citations
- OpenDPR: Open-Vocabulary Change Detection via Vision-Centric Diffusion-Guided Prototype Retrieval for Remote Sensing ImageryQi Guo, Jue Wang, Yinhe Liu, Yanfei ZhongCVPR 2026 · 2 citations
- MARCO: Navigating the Unseen Space of Semantic CorrespondenceClaudia Cuttano, Gabriele Trivigno, Carlo Masone, Stefan RothCVPR 2026 · 4 citations
