A Generative Approach for Wikipedia-Scale Visual Entity Recognition
Mathilde Caron, Ahmet Iscen, Alireza Fathi, Cordelia Schmid
摘要
In this paper, we address web-scale visual entity recognition, specifically the task of mapping a given query image to one of the 6 million existing entities in Wikipedia. One way of approaching a problem of such scale is using dualencoder models (e.g. CLIP), where all the entity names and query images are embedded into a unified space, paving the way for an approximate kNN search. Alternatively, it is also possible to re-purpose a captioning model to directly generate the entity names for a given image. In contrast, we introduce a novel Generative Entity Recognition (GER) framework, which given an input image learns to auto-regressively decode a semantic and discriminative "code" identifying the target entity. Our experiments demonstrate the e cacy of this GER paradigm, showcasing state-of-the-art performance on the challenging OVEN benchmark. GER surpasses strong captioning, dual-encoder, visual matching and hierarchical classification baselines, a rming its advantage in tackling the complexities of web-scale recognition.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- ShotBench: Expert-Level Cinematic Understanding in Vision-Language ModelsHongbo Liu, Jingwen He, Yi Jin, Dian Zheng 等NeurIPS 2025 · 被引用 24 次
- Web-Scale Visual Entity Recognition: An LLM-Driven Data ApproachMathilde Caron, Alireza Fathi, Cordelia Schmid, Ahmet IscenNeurIPS 2024 · 被引用 5 次
- Visual Text Matters: Improving Text-KVQA with Visual Text Entity Knowledge-aware Large Multimodal AssistantAbhirama Subramanyam Penamakuri, Anand MishraEMNLP 2024 · 被引用 2 次
- Seeing and Knowing in the Wild: Open-domain Visual Entity Recognition with Large-scale Knowledge Graphs via Contrastive LearningHongkuan Zhou, Lavdim Halilaj, Sebastian Monka, Stefan Schmid 等AAAI 2026 · 被引用 1 次
- On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action UnderstandingZhanzhong Pang, Dibyadip Chatterjee, Fadime Sener, Angela YaoICLR 2026 · 被引用 1 次
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
- Recommender Systems with Generative RetrievalShashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan 等NeurIPS 2023 · 被引用 474 次
相关 Paper
- Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia EntitiesHexiang Hu, Yi Luan, Yang Chen, Urvashi Khandelwal 等ICCV 2023 · 被引用 123 次
- WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity RecognitionShan Ning, Longtian Qiu, Jiaxuan Sun, Xuming HeCVPR 2026 · 被引用 1 次
- MOFI: Learning Image Representations from Noisy Entity Annotated ImagesWentao Wu, Aleksei Timofeev, Chen Chen, Bowen Zhang 等ICLR 2024 · 被引用 9 次
- Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity LinkingZhengfei Xu, Sijia Zhao, Yanchao Hao, Xiaolong Liu 等AAAI 2025
- Autoregressive Entity RetrievalNicola De Cao, Gautier Izacard, Sebastian Riedel, Fabio PetroniICLR 2021 · 被引用 200 次
