Breaking Common Sense: WHOOPS! A Vision-and-Language Benchmark of Synthetic and Compositional Images
Nitzan Bitton Guetta, Yonatan Bitton, Jack Hessel, Ludwig Schmidt, Yuval Elovici, Gabriel Stanovsky, Roy Schwartz
摘要
Weird, unusual, and uncanny images pique the curiosity of observers because they challenge commonsense. For example, an image released during the 2022 world cup depicts the famous soccer stars Lionel Messi and Cristiano Ronaldo playing chess, which playfully violates our expectation that their competition should occur on the football field. 1 Humans can easily recognize and interpret these unconventional images, but can AI models do the same? We introduce WHOOPS!, a new dataset and benchmark for visual commonsense. The dataset is comprised of purposefully commonsense-defying images created by designers using publicly-available image generation tools like Midjourney. We consider several tasks posed over the dataset. In addition to image captioning, cross-modal matching, and visual question answering, we introduce a difficult explanation generation task, where models must identify and explain why a given image is unusual. Our results show that state-of-the-art models such as GPT3 and BLIP2 still lag behind human performance on WHOOPS!. We hope our dataset will inspire the development of AI models with stronger visual commonsense reasoning abilities. 2
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper42
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 等ICLR 2024 · 被引用 1,472 次
- Grounding Multimodal Large Language Models to the WorldZhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao 等ICLR 2024 · 被引用 1,170 次
- What You See is What You Read? Improving Text-Image Alignment EvaluationMichal Yarom, Yonatan Bitton, Soravit Changpinyo, Roee Aharoni 等NeurIPS 2023 · 被引用 147 次
- Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image GenerationJaemin Cho, Yushi Hu, Jason M. Baldridge, Roopal Garg 等ICLR 2024 · 被引用 139 次
- Vision Language Models are BiasedAn Vo, Khai-Nguyen Nguyen, Mohammad Reza Taesiri, Thi Tuong Vy Dang 等ICLR 2026 · 被引用 68 次
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- FOCUS: Evaluating Pre-trained Vision-Language Models on Underspecification ReasoningKankan Zhou, Eason Lai, Kyriakos Mouratidis, Jing JiangACL 2025
- PhD: A ChatGPT-Prompted Visual Hallucination Evaluation DatasetJiazhen Liu, Yuhan Fu, Ruobing Xie, Runquan Xie 等CVPR 2025
- When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language ModelsFrancesco Ortu, Zhijing Jin, Diego Doimo, Alberto CazzanigaACL 2026 · 被引用 7 次
- YesBut: A High-Quality Annotated Multimodal Dataset for evaluating Satire Comprehension capability of Vision-Language ModelsAbhilash Nandy, Yash Agarwal, Ashish Patwa, Millon Madhur Das 等EMNLP 2024 · 被引用 2 次
- VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video UnderstandingZongxia Li, Xiyang Wu, Guangyao Shi, Yubin Qin 等NeurIPS 2025 · 被引用 38 次
