VASR: Visual Analogies of Situation Recognition
Yonatan Bitton, Ron Yosef, Eliyahu Strugo, Dafna Shahaf, Roy Schwartz, Gabriel Stanovsky
Abstract
A core process in human cognition is analogical mapping: the ability to identify a similar relational structure between different situations. We introduce a novel task, Visual Analogies of Situation Recognition, adapting the classical word-analogy task into the visual domain. Given a triplet of images, the task is to select an image candidate B' that completes the analogy (A to A' is like B to what?). Unlike previous work on visual analogy that focused on simple image transformations, we tackle complex analogies requiring understanding of scenes.
We leverage situation recognition annotations and the CLIP model to generate a large set of 500k candidate analogies. Crowdsourced annotations for a sample of the data indicate that humans agree with the dataset label 80% of the time (chance level 25%). Furthermore, we use human annotations to create a gold-standard dataset of 3,820 validated analogies. Our experiments demonstrate that state-of-the-art models do well when distractors are chosen randomly ( 86%), but struggle with carefully chosen distractors ( 53%, compared to 90% human accuracy). We hope our dataset will encourage the development of new analogy-making models. Website: https://vasr-dataset.github.io/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- Breaking Common Sense: WHOOPS! A Vision-and-Language Benchmark of Synthetic and Compositional ImagesNitzan Bitton Guetta, Yonatan Bitton, Jack Hessel, Ludwig Schmidt et al.ICCV 2023 · 92 citations
- VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain KnowledgeYueqi Song, Tianyue Ou, Yibo Kong, Zecheng Li et al.ICML 2026 · 44 citations
- In-Context Analogical Reasoning with Pre-Trained Language ModelsXiaoyang Hu, Shane Storks, Richard L. Lewis, Joyce ChaiACL 2023 · 13 citations
- StoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical UnderstandingCheng Jiayang, Lin Qiu, Tsz Ho Chan, Tianqing Fang et al.EMNLP 2023 · 8 citations
- Bongard-RWR+: Real-World Representations of Fine-Grained Concepts in Bongard ProblemsSzymon Pawlonka, Mikołaj Małkiński, Jacek MańdziukICLR 2026 · 7 citations
Builds on6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
Related papers
- Relational Visual SimilarityThao Nguyen, Sicheng Mo, Krishna Kumar Singh, Yilin Wang et al.CVPR 2026 · 1 citation
- Why Does a Visual Question Have Different Answers?Nilavra Bhattacharya, Qing Li, Danna GurariICCV 2019 · 78 citations
- Multimodal Analogical Reasoning over Knowledge GraphsNingyu Zhang, Lei Li, Xiang Chen, Xiaozhuan Liang et al.ICLR 2023 · 10 citations
- Few-shot Visual Reasoning with Meta-Analogical Contrastive LearningYoungsung Kim, Jinwoo Shin, Eunho Yang, Sung Ju HwangNeurIPS 2020 · 30 citations
- MemeCap: A Dataset for Captioning and Interpreting MemesEunjeong Hwang, Vered ShwartzEMNLP 2023 · 14 citations
