Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
Mathilde Caron, Alireza Fathi, Cordelia Schmid, Ahmet Iscen
摘要
Web-scale visual entity recognition, the task of associating images with their corresponding entities within vast knowledge bases like Wikipedia, presents significant challenges due to the lack of clean, large-scale training data. In this paper, we propose a novel methodology to curate such a dataset, leveraging a multimodal large language model (LLM) for label verification, metadata generation, and rationale explanation. Instead of relying on the multimodal LLM to directly annotate data, which we found to be suboptimal, we prompt it to reason about potential candidate entity labels by accessing additional contextually relevant information (such as Wikipedia), resulting in more accurate annotations. We further use the multimodal LLM to enrich the dataset by generating question-answer pairs and a grounded finegrained textual description (referred to as"rationale") that explains the connection between images and their assigned entities. Experiments demonstrate that models trained on this automatically curated data achieve state-of-the-art performance on web-scale visual entity recognition tasks (e.g. +6.9% improvement in OVEN entity task), underscoring the importance of high-quality training data in this domain.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ShotBench: Expert-Level Cinematic Understanding in Vision-Language ModelsHongbo Liu, Jingwen He, Yi Jin, Dian Zheng 等NeurIPS 2025 · 被引用 24 次
- BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation ModelsShengao Wang, Wenqi Wang, Zecheng Wang, Max Whitton 等CVPR 2026 · 被引用 4 次
- WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity RecognitionShan Ning, Longtian Qiu, Jiaxuan Sun, Xuming HeCVPR 2026 · 被引用 1 次
- How to Take a Memorable Picture? Empowering Users with Actionable FeedbackFrancesco Laiti, Davide Talon, Jacopo Staiano, Elisa RicciCVPR 2026
- Stabilizing Feature Geometry in Noisy Pretrained Models for Robust Downstream TasksQuanyu Zhang, Zhongyi Han, Hao Sun, Yongshun Gong 等CVPR 2026
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Transformer Memory as a Differentiable Search IndexYi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni 等NeurIPS 2022 · 被引用 506 次
- Recommender Systems with Generative RetrievalShashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan 等NeurIPS 2023 · 被引用 474 次
- Just Ask: Learning to Answer Questions from Millions of Narrated VideosAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 等ICCV 2021 · 被引用 345 次
相关 Paper
- Large Language Models Know What is Key Visual Entity: An LLM-assisted Multimodal Retrieval for VQAPu Jian, Donglei Yu, Jiajun ZhangEMNLP 2024 · 被引用 5 次
- A Generative Approach for Wikipedia-Scale Visual Entity RecognitionMathilde Caron, Ahmet Iscen, Alireza Fathi, Cordelia SchmidCVPR 2024
- Grounding Multilingual Multimodal LLMs With Cultural KnowledgeJean de Dieu Nyandwi, Yueqi Song, Simran Khanuja, Graham NeubigEMNLP 2025
- Multimodal Entity Linking: A New Dataset and A BaselineJingru Gan, Jinchang Luo, Haiwei Wang, Shuhui Wang 等ACM MM 2021 · 被引用 41 次
- VisualWebInstruct: Scaling up Multimodal Instruction Data through Web SearchYiming Jia, Jiachen Li, Xiang Yue, Bo Li 等EMNLP 2025 · 被引用 29 次
