When Pigs Fly: Contextual Reasoning in Synthetic and Natural Scenes
Philipp Bomatter, Mengmi Zhang, Dimitar Karev, Spandan Madan, Claire Tseng, Gabriel Kreiman
摘要
Context is of fundamental importance to both human and machine vision; e.g., an object in the air is more likely to be an airplane than a pig. The rich notion of context incorporates several aspects including physics rules, statistical co-occurrences, and relative object sizes, among others. While previous work has focused on crowd-sourced out-of-context photographs from the web to study scene context, controlling the nature and extent of contextual violations has been a daunting task. Here we introduce a diverse, synthetic Out-of-Context Dataset (OCD) with fine-grained control over scene context. By leveraging a 3D simulation engine, we systematically control the gravity, object co-occurrences and relative sizes across 36 object categories in a virtual household environment. We conducted a series of experiments to gain insights into the impact of contextual cues on both human and machine vision using OCD. We conducted psychophysics experiments to establish a human benchmark for out-of-context recognition, and then compared it with state-of-the-art computer vision models to quantify the gap between the two. We propose a context-aware recognition transformer model, fusing object and contextual information via multi-head attention. Our model captures useful information for contextual reasoning, enabling human-level performance and better robustness in out-of-context conditions compared to baseline models across OCD and other out-of-context datasets. All source code and data are publicly available at https://github.com/kreimanlab/WhenPigsFlyContext
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Adaptive Visual Scene Understanding: Incremental Scene Graph GenerationNaitik Khandelwal, Xiao Liu, Mengmi ZhangNeurIPS 2024 · 被引用 9 次
- Visual Data Diagnosis and Debiasing with Concept GraphsRwiddhi Chakraborty, Yinong Wang, Jialu Gao, Runkai Zheng 等NeurIPS 2024 · 被引用 9 次
- Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality FusionIshaan Singh Rawal, Alexander Matyasko, Shantanu Jaiswal, Basura Fernando 等ICML 2024 · 被引用 8 次
- Common Inpainted Objects In-N-Out of ContextTianze Yang, Tyson Jordan, Ruitong Sun, Ninghao Liu 等CVPR 2026 · 被引用 8 次
- Learning to See Through a Baby’s Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and MachinesYusen Cai, Qing Lin, BHARGAVA SATYA NUNNA, Mengmi ZhangCVPR 2026 · 被引用 4 次
它引用的顶会 Paper5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Physics-Based Rendering for Improving Robustness to RainShirsendu Sukanta Halder, Jean-François Lalonde, Raoul de CharetteICCV 2019 · 被引用 129 次
- Putting Visual Object Recognition in ContextMengmi Zhang, Claire Tseng, Gabriel KreimanCVPR 2020
- Simple Copy-Paste Is a Strong Data Augmentation Method for Instance SegmentationGolnaz Ghiasi, Yin Cui, Aravind Srinivas, Rui Qian 等CVPR 2021
- Learning From Synthetic AnimalsJiteng Mu, Weichao Qiu, Gregory D. Hager, Alan L. YuilleCVPR 2020
相关 Paper
- Scene Context-Aware Salient Object DetectionAvishek Siris, Jianbo Jiao, Gary K. L. Tam, Xianghua Xie 等ICCV 2021 · 被引用 54 次
- ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language ModelsZhaoyang Li, Zhan Ling, Yuchen Zhou, Litian Gong 等CVPR 2026 · 被引用 2 次
- ORIDa: Object-centric Real-world Image Composition DatasetJinwoo Kim, Sangmin Han, Jinho Jeong, Jiwoo Choi 等CVPR 2025
- CIRCLE: Capture In Rich Contextual EnvironmentsJoão Pedro Araújo, Jiaman Li, Karthik Vetrivel, Rishi Agarwal 等CVPR 2023
- Adaptive Contextual Perception: How To Generalize To New Backgrounds and Ambiguous ObjectsZhuofan Ying, Peter Hase, Mohit BansalNeurIPS 2023 · 被引用 2 次
