VASR: Visual Analogies of Situation Recognition
Yonatan Bitton, Ron Yosef, Eliyahu Strugo, Dafna Shahaf, Roy Schwartz, Gabriel Stanovsky
摘要
A core process in human cognition is analogical mapping: the ability to identify a similar relational structure between different situations. We introduce a novel task, Visual Analogies of Situation Recognition, adapting the classical word-analogy task into the visual domain. Given a triplet of images, the task is to select an image candidate B' that completes the analogy (A to A' is like B to what?). Unlike previous work on visual analogy that focused on simple image transformations, we tackle complex analogies requiring understanding of scenes.
We leverage situation recognition annotations and the CLIP model to generate a large set of 500k candidate analogies. Crowdsourced annotations for a sample of the data indicate that humans agree with the dataset label 80% of the time (chance level 25%). Furthermore, we use human annotations to create a gold-standard dataset of 3,820 validated analogies. Our experiments demonstrate that state-of-the-art models do well when distractors are chosen randomly ( 86%), but struggle with carefully chosen distractors ( 53%, compared to 90% human accuracy). We hope our dataset will encourage the development of new analogy-making models. Website: https://vasr-dataset.github.io/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Breaking Common Sense: WHOOPS! A Vision-and-Language Benchmark of Synthetic and Compositional ImagesNitzan Bitton Guetta, Yonatan Bitton, Jack Hessel, Ludwig Schmidt 等ICCV 2023 · 被引用 92 次
- VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain KnowledgeYueqi Song, Tianyue Ou, Yibo Kong, Zecheng Li 等ICML 2026 · 被引用 44 次
- In-Context Analogical Reasoning with Pre-Trained Language ModelsXiaoyang Hu, Shane Storks, Richard L. Lewis, Joyce ChaiACL 2023 · 被引用 13 次
- StoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical UnderstandingCheng Jiayang, Lin Qiu, Tsz Ho Chan, Tianqing Fang 等EMNLP 2023 · 被引用 8 次
- Bongard-RWR+: Real-World Representations of Fine-Grained Concepts in Bongard ProblemsSzymon Pawlonka, Mikołaj Małkiński, Jacek MańdziukICLR 2026 · 被引用 7 次
它引用的顶会 Paper6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
相关 Paper
- Relational Visual SimilarityThao Nguyen, Sicheng Mo, Krishna Kumar Singh, Yilin Wang 等CVPR 2026 · 被引用 1 次
- Why Does a Visual Question Have Different Answers?Nilavra Bhattacharya, Qing Li, Danna GurariICCV 2019 · 被引用 78 次
- Multimodal Analogical Reasoning over Knowledge GraphsNingyu Zhang, Lei Li, Xiang Chen, Xiaozhuan Liang 等ICLR 2023 · 被引用 10 次
- Few-shot Visual Reasoning with Meta-Analogical Contrastive LearningYoungsung Kim, Jinwoo Shin, Eunho Yang, Sung Ju HwangNeurIPS 2020 · 被引用 30 次
- MemeCap: A Dataset for Captioning and Interpreting MemesEunjeong Hwang, Vered ShwartzEMNLP 2023 · 被引用 14 次
