When Pigs Fly: Contextual Reasoning in Synthetic and Natural Scenes
Philipp Bomatter, Mengmi Zhang, Dimitar Karev, Spandan Madan, Claire Tseng, Gabriel Kreiman
Abstract
Context is of fundamental importance to both human and machine vision; e.g., an object in the air is more likely to be an airplane than a pig. The rich notion of context incorporates several aspects including physics rules, statistical co-occurrences, and relative object sizes, among others. While previous work has focused on crowd-sourced out-of-context photographs from the web to study scene context, controlling the nature and extent of contextual violations has been a daunting task. Here we introduce a diverse, synthetic Out-of-Context Dataset (OCD) with fine-grained control over scene context. By leveraging a 3D simulation engine, we systematically control the gravity, object co-occurrences and relative sizes across 36 object categories in a virtual household environment. We conducted a series of experiments to gain insights into the impact of contextual cues on both human and machine vision using OCD. We conducted psychophysics experiments to establish a human benchmark for out-of-context recognition, and then compared it with state-of-the-art computer vision models to quantify the gap between the two. We propose a context-aware recognition transformer model, fusing object and contextual information via multi-head attention. Our model captures useful information for contextual reasoning, enabling human-level performance and better robustness in out-of-context conditions compared to baseline models across OCD and other out-of-context datasets. All source code and data are publicly available at https://github.com/kreimanlab/WhenPigsFlyContext
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3703ebec-25de-47f7-b22d-cc9099353b6eCited by top-tier papers7
- Adaptive Visual Scene Understanding: Incremental Scene Graph GenerationNaitik Khandelwal, Xiao Liu, Mengmi ZhangNeurIPS 2024 · 9 citations
- Visual Data Diagnosis and Debiasing with Concept GraphsRwiddhi Chakraborty, Yinong Wang, Jialu Gao, Runkai Zheng et al.NeurIPS 2024 · 9 citations
- Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality FusionIshaan Singh Rawal, Alexander Matyasko, Shantanu Jaiswal, Basura Fernando et al.ICML 2024 · 8 citations
- Common Inpainted Objects In-N-Out of ContextTianze Yang, Tyson Jordan, Ruitong Sun, Ninghao Liu et al.CVPR 2026 · 8 citations
- Learning to See Through a Baby’s Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and MachinesYusen Cai, Qing Lin, BHARGAVA SATYA NUNNA, Mengmi ZhangCVPR 2026 · 4 citations
Builds on5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Physics-Based Rendering for Improving Robustness to RainShirsendu Sukanta Halder, Jean-François Lalonde, Raoul de CharetteICCV 2019 · 129 citations
- Putting Visual Object Recognition in ContextMengmi Zhang, Claire Tseng, Gabriel KreimanCVPR 2020
- Simple Copy-Paste Is a Strong Data Augmentation Method for Instance SegmentationGolnaz Ghiasi, Yin Cui, Aravind Srinivas, Rui Qian et al.CVPR 2021
- Learning From Synthetic AnimalsJiteng Mu, Weichao Qiu, Gregory D. Hager, Alan L. YuilleCVPR 2020
Related papers
- Scene Context-Aware Salient Object DetectionAvishek Siris, Jianbo Jiao, Gary K. L. Tam, Xianghua Xie et al.ICCV 2021 · 54 citations
- ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language ModelsZhaoyang Li, Zhan Ling, Yuchen Zhou, Litian Gong et al.CVPR 2026 · 2 citations
- ORIDa: Object-centric Real-world Image Composition DatasetJinwoo Kim, Sangmin Han, Jinho Jeong, Jiwoo Choi et al.CVPR 2025
- CIRCLE: Capture In Rich Contextual EnvironmentsJoão Pedro Araújo, Jiaman Li, Karthik Vetrivel, Rishi Agarwal et al.CVPR 2023
- Adaptive Contextual Perception: How To Generalize To New Backgrounds and Ambiguous ObjectsZhuofan Ying, Peter Hase, Mohit BansalNeurIPS 2023 · 2 citations
