UDON: Universal Dynamic Online distillatioN for generic image representations
Nikolaos-Antonios Ypsilantis, Kaifeng Chen, André Araújo, Ondrej Chum
Abstract
Universal image representations are critical in enabling real-world fine-grained and instance-level recognition applications, where objects and entities from any domain must be identified at large scale. Despite recent advances, existing methods fail to capture important domain-specific knowledge, while also ignoring differences in data distribution across different domains. This leads to a large performance gap between efficient universal solutions and expensive approaches utilising a collection of specialist models, one for each domain. In this work, we make significant strides towards closing this gap, by introducing a new learning technique, dubbed UDON (Universal Dynamic Online DistillatioN). UDON employs multi-teacher distillation, where each teacher is specialized in one domain, to transfer detailed domain-specific knowledge into the student universal embedding. UDON's distillation approach is not only effective, but also very efficient, by sharing most model parameters between the student and all teachers, where all models are jointly trained in an online manner. UDON also comprises a sampling technique which adapts the training process to dynamically allocate batches to domains which are learned slower and require more frequent processing. This boosts significantly the learning of complex domains which are characterised by a large number of classes and long-tail distributions. With comprehensive experiments, we validate each component of UDON, and showcase significant improvements over the state of the art in the recent UnED benchmark. Code: https://github.com/nikosips/UDON .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45d64473-be39-43b5-a9db-0f5552ed20bdCited by top-tier papers5
- Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMsZhiyu Pan, Yizheng Wu, Jiashen Hua, Junyi Feng et al.ICLR 2026 · 11 citations
- Illuminating Visual Identity in Universal Multimodal EmbeddingsJiawei Cao, Junyi Feng, Jiashen Hua, Ziheng Huang et al.CVPR 2026 · 1 citation
- DUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D TeachersMert Bülent Sariyildiz, Philippe Weinzaepfel, Thomas Lucas, Pau de Jorge et al.CVPR 2025
- Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation ModelsJiawei Fan, Shigeng Wang, Chao Li, Xiaolong Liu et al.CVPR 2026
- ILIAS: Instance-Level Image retrieval At ScaleGiorgos Kordopatis-Zilos, Vladan Stojnic, Anna Manko, Pavel Suma et al.CVPR 2025
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Hyperbolic Vision Transformers: Combining Improvements in Metric LearningAleksandr Ermolov, Leyla Mirvakhabova, Valentin Khrulkov, Nicu Sebe et al.CVPR 2022 · 97 citations
- Recall@k Surrogate Loss with Large Batches and Similarity MixupYash Patel, Giorgos Tolias, Jirí MatasCVPR 2022 · 40 citations
Related papers
- DomED: Redesigning Ensemble Distillation for Domain GeneralizationZiang Song, Zhou Zhidan, Zijun ZhangICML 2026
- Distilling from Similar Tasks for Transfer Learning on a BudgetKenneth Borup, Cheng Perng Phoo, Bharath HariharanICCV 2023 · 3 citations
- Towards Universal Image Embeddings: A Large-Scale Dataset and Challenge for Generic Image RepresentationsNikolaos-Antonios Ypsilantis, Kaifeng Chen, Bingyi Cao, Mário Lipovský et al.ICCV 2023 · 31 citations
- Unsupervised Video Domain Adaptation with Masked Pre-Training and Collaborative Self-TrainingArun V. Reddy, William Paul, Corban Rivera, Ketul Shah et al.CVPR 2024 · 3 citations
- Multi-Target Adversarial Frameworks for Domain Adaptation in Semantic SegmentationAntoine Saporta, Tuan-Hung Vu, Matthieu Cord, Patrick PérezICCV 2021 · 41 citations
