Putting Visual Object Recognition in Context
Mengmi Zhang, Claire Tseng, Gabriel Kreiman
Abstract
Context plays an important role in visual recognition. Recent studies have shown that visual recognition networks can be fooled by placing objects in inconsistent contexts (e.g., a cow in the ocean). To model the role of contextual information in visual recognition, we systematically investigated ten critical properties of where, when, and how context modulates recognition, including the amount of context, context and object resolution, geometrical structure of context, context congruence, and temporal dynamics of contextual modulation. The tasks involved recognizing a target object surrounded with context in a natural image. As an essential benchmark, we conducted a series of psychophysics experiments where we altered one aspect of context at a time, and quantified recognition accuracy. We propose a biologically-inspired context-aware object recognition model consisting of a two-stream architecture. The model processes visual information at the fovea and periphery in parallel, dynamically incorporates object and contextual information, and sequentially reasons about the class label for the target object. Across a wide range of behavioral tasks, the model approximates human level performance without retraining for each task, captures the dependence of context enhancement on image properties, and provides initial steps towards integrating scene and object information for visual recognition. All source code and data are publicly available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2226b9d4-5cb3-44b7-8e9a-9471bc9d21cfCited by top-tier papers8
- Discriminative Region-based Multi-Label Zero-Shot LearningSanath Narayan, Akshita Gupta, Salman H. Khan, Fahad Shahbaz Khan et al.ICCV 2021 · 62 citations
- When Pigs Fly: Contextual Reasoning in Synthetic and Natural ScenesPhilipp Bomatter, Mengmi Zhang, Dimitar Karev, Spandan Madan et al.ICCV 2021 · 30 citations
- Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality FusionIshaan Singh Rawal, Alexander Matyasko, Shantanu Jaiswal, Basura Fernando et al.ICML 2024 · 8 citations
- Common Inpainted Objects In-N-Out of ContextTianze Yang, Tyson Jordan, Ruitong Sun, Ninghao Liu et al.CVPR 2026 · 8 citations
- Learning to See Through a Baby’s Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and MachinesYusen Cai, Qing Lin, BHARGAVA SATYA NUNNA, Mengmi ZhangCVPR 2026 · 4 citations
Related papers
- Flexible Context-Driven Sensory Processing in Dynamical Vision ModelsLakshmi Narasimhan Govindarajan, Abhiram Iyer, Valmiki Kothare, Ila FieteNeurIPS 2024 · 1 citation
- Don't Judge an Object by Its Context: Learning to Overcome Contextual BiasKrishna Kumar Singh, Dhruv Mahajan, Kristen Grauman, Yong Jae Lee et al.CVPR 2020
- Uni-Perceiver: Pre-training Unified Architecture for Generic Perception for Zero-shot and Few-shot TasksXizhou Zhu, Jinguo Zhu, Hao Li, Xiaoshi Wu et al.CVPR 2022
- Peripheral Vision TransformerJuhong Min, Yucheng Zhao, Chong Luo, Minsu ChoNeurIPS 2022 · 49 citations
- Adaptive Contextual Perception: How To Generalize To New Backgrounds and Ambiguous ObjectsZhuofan Ying, Peter Hase, Mohit BansalNeurIPS 2023 · 2 citations
