Who's Waldo? Linking People Across Text and Images
Claire Yuqing Cui, Apoorv Khandelwal, Yoav Artzi, Noah Snavely, Hadar Averbuch-Elor
摘要
We present a task and benchmark dataset for person-centric visual grounding, the problem of linking between people named in a caption and people pictured in an image. In contrast to prior work in visual grounding, which is predominantly object-based, our new task masks out the names of people in captions in order to encourage methods trained on such image–caption pairs to focus on contextual cues, such as the rich interactions between multiple people, rather than learning associations between names and appearances. To facilitate this task, we introduce a new dataset, Who’s Waldo, mined automatically from image–caption data on Wikimedia Commons. We propose a Transformer-based method that outperforms several strong baselines on this task, and release our data to the research community to spur work on contextual models that consider both vision and language. Code and data are available at: https://whoswaldo.github.io
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Who are you referring to? Coreference resolution in image narrationsArushi Goel, Basura Fernando, Frank Keller, Hakan BilenICCV 2023 · 被引用 8 次
- Learning Human-Human Interactions in Images from Weak Textual SupervisionMorris Alper, Hadar Averbuch-ElorICCV 2023 · 被引用 4 次
- Semi-supervised multimodal coreference resolution in image narrationsArushi Goel, Basura Fernando, Frank Keller, Hakan BilenEMNLP 2023 · 被引用 4 次
- There's a Time and Place for Reasoning Beyond the ImageXingyu Fu, Ben Zhou, Ishaan Preetam Chandratreya, Carl Vondrick 等ACL 2022
- JRDB-Social: A Multifaceted Robotic Dataset for Understanding of Context and Dynamics of Human Interactions Within Social GroupsSimindokht Jahangard, Zhixi Cai, Shiki Wen, Hamid RezatofighiCVPR 2024
它引用的顶会 Paper10
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu 等AAAI 2020 · 被引用 1,047 次
- Large-Scale Adversarial Training for Vision-and-Language Representation LearningZhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu 等NeurIPS 2020 · 被引用 561 次
- Learning to Assemble Neural Module Tree Networks for Visual GroundingDaqing Liu, Hanwang Zhang, Feng Wu, Zheng-Jun ZhaICCV 2019 · 被引用 317 次
相关 Paper
- Linking People across Text and Images Based on Social Relation ReasoningYang Lei, Peizhi Zhao, Pijian Li, Yi Cai 等AAAI 2023
- MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMsYunqiu Xu, Linchao Zhu, Yi YangICCV 2025 · 被引用 7 次
- Detecting and Grounding Important Characters in Visual StoriesDanyang Liu, Frank KellerAAAI 2023 · 被引用 11 次
- AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional RelationsJunli Liu, Qizhi Chen, Zhigang Wang, Yiwen Tang 等ICCV 2025 · 被引用 5 次
- When Visual Grounding Meets Gigapixel-Level Large-Scale Scenes: Benchmark and ApproachM. Tao, Bing Bai, Haozhe Lin, Heyuan Wang 等CVPR 2024 · 被引用 4 次
