Towards Universal Image Embeddings: A Large-Scale Dataset and Challenge for Generic Image Representations
Nikolaos-Antonios Ypsilantis, Kaifeng Chen, Bingyi Cao, Mário Lipovský, Pelin Dogan-Schönberger, Grzegorz Makosa, Boris Bluntschli, Mojtaba Seyedhosseini, Ondrej Chum, André Araújo
摘要
Fine-grained and instance-level recognition methods are commonly trained and evaluated on specific domains, in a model per domain scenario. Such an approach, however, is impractical in real large-scale applications. In this work, we address the problem of universal image embedding, where a single universal model is trained and used in multiple domains. First, we leverage existing domain-specific datasets to carefully construct a new large-scale public benchmark for the evaluation of universal image embeddings, with 241k query images, 1.4M index images and 2.8M training images across 8 different domains and 349k classes. We define suitable metrics, training and evaluation protocols to foster future research in this area. Second, we provide a comprehensive experimental evaluation on the new dataset, demonstrating that existing approaches and simplistic extensions lead to worse performance than an assembly of models trained for each domain separately. Finally, we conducted a public research competition on this topic, leveraging industrial datasets, which attracted the participation of more than 1k teams world-wide. This exercise generated many interesting research ideas and findings which we present in detail. Project webpage: https://cmp.felk.cvut.cz/univ_emb/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Residual Quantization with Implicit Neural CodebooksIris A. M. Huijben, Matthijs Douze, Matthew J. Muckley, Ruud van Sloun 等ICML 2024 · 被引用 23 次
- TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text AlignmentBingyi Cao, Koert Chen, Kevis-Kokitsi Maninis, Kaifeng Chen 等CVPR 2026 · 被引用 14 次
- Instance-Level Composed Image RetrievalBill Psomas, George Retsinas, Nikos Efthymiadis, Panagiotis Paraskevas Filntisis 等NeurIPS 2025 · 被引用 14 次
- UDON: Universal Dynamic Online distillatioN for generic image representationsNikolaos-Antonios Ypsilantis, Kaifeng Chen, André Araújo, Ondrej ChumNeurIPS 2024 · 被引用 12 次
- Threshold-Consistent Margin Loss for Open-World Deep Metric LearningQin Zhang, Linghan Xu, Jun Fang, Qingming Tang 等ICLR 2024 · 被引用 10 次
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- AdaFace: Quality Adaptive Margin for Face RecognitionMinchul Kim, Anil K. Jain, Xiaoming LiuCVPR 2022 · 被引用 509 次
- DOLG: Single-Stage Image Retrieval with Deep Orthogonal Fusion of Local and Global FeaturesMin Yang, Dongliang He, Miao Fan, Baorong Shi 等ICCV 2021 · 被引用 135 次
相关 Paper
- Google Landmarks Dataset v2 - A Large-Scale Benchmark for Instance-Level Recognition and RetrievalTobias Weyand, André Araújo, Bingyi Cao, Jack SimCVPR 2020
- Adaptive Methods for Real-World Domain GeneralizationAbhimanyu Dubey, Vignesh Ramanathan, Alex Pentland, Dhruv MahajanCVPR 2021
- Mieb: Massive Image Embedding BenchmarkChenghao Xiao, Isaac Chung, Imene Kerboua, Jamie Stirling 等ICCV 2025
- Illuminating Visual Identity in Universal Multimodal EmbeddingsJiawei Cao, Junyi Feng, Jiashen Hua, Ziheng Huang 等CVPR 2026 · 被引用 1 次
- VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding TasksZiyan Jiang, Rui Meng, Xinyi Yang, Semih Yavuz 等ICLR 2025
