A Generative Approach for Wikipedia-Scale Visual Entity Recognition
Mathilde Caron, Ahmet Iscen, Alireza Fathi, Cordelia Schmid
Abstract
In this paper, we address web-scale visual entity recognition, specifically the task of mapping a given query image to one of the 6 million existing entities in Wikipedia. One way of approaching a problem of such scale is using dualencoder models (e.g. CLIP), where all the entity names and query images are embedded into a unified space, paving the way for an approximate kNN search. Alternatively, it is also possible to re-purpose a captioning model to directly generate the entity names for a given image. In contrast, we introduce a novel Generative Entity Recognition (GER) framework, which given an input image learns to auto-regressively decode a semantic and discriminative "code" identifying the target entity. Our experiments demonstrate the e cacy of this GER paradigm, showcasing state-of-the-art performance on the challenging OVEN benchmark. GER surpasses strong captioning, dual-encoder, visual matching and hierarchical classification baselines, a rming its advantage in tackling the complexities of web-scale recognition.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9697815b-99db-49d3-8e5b-9b79ecd300bdCited by top-tier papers9
- ShotBench: Expert-Level Cinematic Understanding in Vision-Language ModelsHongbo Liu, Jingwen He, Yi Jin, Dian Zheng et al.NeurIPS 2025 · 24 citations
- Web-Scale Visual Entity Recognition: An LLM-Driven Data ApproachMathilde Caron, Alireza Fathi, Cordelia Schmid, Ahmet IscenNeurIPS 2024 · 5 citations
- Visual Text Matters: Improving Text-KVQA with Visual Text Entity Knowledge-aware Large Multimodal AssistantAbhirama Subramanyam Penamakuri, Anand MishraEMNLP 2024 · 2 citations
- Seeing and Knowing in the Wild: Open-domain Visual Entity Recognition with Large-scale Knowledge Graphs via Contrastive LearningHongkuan Zhou, Lavdim Halilaj, Sebastian Monka, Stefan Schmid et al.AAAI 2026 · 1 citation
- On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action UnderstandingZhanzhong Pang, Dibyadip Chatterjee, Fadime Sener, Angela YaoICLR 2026 · 1 citation
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- Recommender Systems with Generative RetrievalShashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan et al.NeurIPS 2023 · 474 citations
Related papers
- Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia EntitiesHexiang Hu, Yi Luan, Yang Chen, Urvashi Khandelwal et al.ICCV 2023 · 123 citations
- WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity RecognitionShan Ning, Longtian Qiu, Jiaxuan Sun, Xuming HeCVPR 2026 · 1 citation
- MOFI: Learning Image Representations from Noisy Entity Annotated ImagesWentao Wu, Aleksei Timofeev, Chen Chen, Bowen Zhang et al.ICLR 2024 · 9 citations
- Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity LinkingZhengfei Xu, Sijia Zhao, Yanchao Hao, Xiaolong Liu et al.AAAI 2025
- Autoregressive Entity RetrievalNicola De Cao, Gautier Izacard, Sebastian Riedel, Fabio PetroniICLR 2021 · 200 citations
