Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia Entities
Hexiang Hu, Yi Luan, Yang Chen, Urvashi Khandelwal, Mandar Joshi, Kenton Lee, Kristina Toutanova, Ming-Wei Chang
Abstract
Large-scale multi-modal pre-training models such as CLIP [30] and PaLI [8] exhibit strong generalization on various visual domains and tasks. However, existing image classification benchmarks often evaluate recognition on a specific domain (e.g., outdoor images) or a specific task (e.g., classifying plant species), which falls short of evaluating whether pre-trained foundational models are universal visual recognizers. To address this, we formally present the task of Open-domain Visual Entity recognitioN (Oven), where a model need to link an image onto a Wikipedia entity with respect to a text query. We construct Oven-Wiki ‡ by repurposing 14 existing datasets with all labels grounded onto one single label space: Wikipedia entities. Oven-Wiki challenges models to select among six million possible Wikipedia entities, making it a general visual recognition benchmark with the largest number of labels. Our study on state-ofthe-art pre-trained models reveals large headroom in generalizing to the massive-scale label space. We show that a PaLI-based auto-regressive visual recognition model performs surprisingly well, even on Wikipedia entities that have never been seen during fine-tuning. We also find existing pretrained models yield different strengths: while PaLI-based models obtain higher overall performance, CLIP-based models are better at recognizing tail entities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e86c068b-2c8c-47f4-b6a7-a8ea363d9316Cited by top-tier papers63
- MagicLens: Self-Supervised Image Retrieval with Open-Ended InstructionsKai Zhang, Yi Luan, Hexiang Hu, Kenton Lee et al.ICML 2024 · 112 citations
- Reinforced Adaptive Knowledge Learning for Multimodal Fake News DetectionLitian Zhang, Xiaoming Zhang, Ziyi Zhou, Feiran Huang et al.AAAI 2024 · 54 citations
- Retrieval-Enhanced Contrastive Vision-Text ModelsAhmet Iscen, Mathilde Caron, Alireza Fathi, Cordelia SchmidICLR 2024 · 44 citations
- MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late InteractionZilin Xiao, Qi Ma, Mengting Gu, Chun-cheng Jason Chen et al.ICLR 2026 · 40 citations
- Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?Yang Chen, Hexiang Hu, Yi Luan, Haitian Sun et al.EMNLP 2023 · 37 citations
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- SimVLM: Simple Visual Language Model Pretraining with Weak SupervisionZirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai et al.ICLR 2022 · 950 citations
Related papers
- A Generative Approach for Wikipedia-Scale Visual Entity RecognitionMathilde Caron, Ahmet Iscen, Alireza Fathi, Cordelia SchmidCVPR 2024
- WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity RecognitionShan Ning, Longtian Qiu, Jiaxuan Sun, Xuming HeCVPR 2026 · 1 citation
- Seeing and Knowing in the Wild: Open-domain Visual Entity Recognition with Large-scale Knowledge Graphs via Contrastive LearningHongkuan Zhou, Lavdim Halilaj, Sebastian Monka, Stefan Schmid et al.AAAI 2026 · 1 citation
- Web-Scale Visual Entity Recognition: An LLM-Driven Data ApproachMathilde Caron, Alireza Fathi, Cordelia Schmid, Ahmet IscenNeurIPS 2024 · 5 citations
- Concept-pedia: a Wide-coverage Semantically-annotated Multimodal DatasetKarim Ghonim, Andrei Stefan Bejgu, Alberte Fernández-Castro, Roberto NavigliEMNLP 2025
