Seeing More with Less: Human-like Representations in Vision Models
Andrey Gizdov, Shimon Ullman, Daniel Harari
Abstract
Large multimodal models (LMMs) typically process visual inputs with uniform resolution across the entire field of view, leading to inefficiencies when non-critical image regions are processed as precisely as key areas. Inspired by the human visual system's foveated approach, we apply a sampling method to leading architectures such as MDETR, BLIP2, InstructBLIP, LLaVA, and ViLT, and evaluate their performance with variable (foveated) resolution inputs. Results show that foveated sampling boosts accuracy in visual tasks like question answering and object detection under tight pixel budgets, improving performance by up to 2.7% on the GQA dataset, 2.1% on SEED-Bench, and 2.0% on VQAv2 compared to uniform sampling. Furthermore, we show that indiscriminate resolution increases yield diminishing returns, with models achieving up to 80% of their full capability using just 3% of the pixels, even on complex tasks. Foveated sampling prompts more human-like processing within models, such as neuronal selectivity and globally acting self-attention in vision transformers. This paper provides a foundational analysis of foveated sampling's impact on existing models, suggesting that more efficient architectural adaptations, mimicking human visual processing, are a promising research venue for the community. Potential applications of our findings center low power minimal bandwidth devices (such as UAVs and edge devices), where compact and efficient vision is critical.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dbb515c3-7d09-4cf3-9c47-5d20eac5d6d8Cited by top-tier papers3
- Through the Magnifying Glass: Adaptive Perception Magnification for Hallucination-Free VLM DecodingShunqi Mao, Chaoyi Zhang, Tom Weidong CaiACL 2026 · 10 citations
- LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language ModelsSoumyaratna Debnath, Bui Manh Duc, Zinan Liu, Lin WangCVPR 2026 · 5 citations
- One Token, Two Fates: A Unified Framework via Vision Token Manipulation Against MLLMs HallucinationZhan Fa, Yue Duan, Jian Zhang, Lei Qi et al.CVPR 2026 · 2 citations
Builds on9
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
Related papers
- LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMsHaoran Lou, Chunxiao Fan, Ziyan Liu, Yuexin Wu et al.ICCV 2025 · 1 citation
- Improved Baselines with Visual Instruction TuningHaotian Liu, Chunyuan Li, Yuheng Li, Yong Jae LeeCVPR 2024
- Visual Compositional TuningXindi Wu, Hee Seung Hwang, Polina Kirichenko, Esin Tureci et al.ICLR 2026 · 3 citations
- HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language ModelsKazi Hasan Ibn Arif, JinYi Yoon, Dimitrios S. Nikolopoulos, Hans Vandierendonck et al.AAAI 2025 · 49 citations
- FOVI: A biologically-inspired foveated interface for deep vision modelsNicholas Blauch, George Alvarez, Talia KonkleICML 2026
