FOVI: A biologically-inspired foveated interface for deep vision models
Nicholas Blauch, George Alvarez, Talia Konkle
Abstract
Human vision is foveated, with variable resolution peaking at the center of a large field of view; this reflects an efficient trade-off for active sensing, allowing eye-movements to bring different parts of the world into focus with other parts of the world in context. In contrast, most computer vision systems encode the visual world at a uniform resolution, raising challenges for processing full-field high-resolution images efficiently. We propose a foveated vision interface (FOVI) based on the human retina and primary visual cortex (V1), that reformats a variable-resolution retina-like sensor array into a uniformly dense, V1-like sensor manifold. Receptive fields are defined as k-nearest-neighborhoods (kNNs) on the sensor manifold, enabling kNN-convolution via a novel kernel mapping technique. We demonstrate two use cases: (1) an end-to-end kNN-convolutional architecture, and (2) a foveated adaptation of the DINOv3 ViT foundation model, leveraging low-rank adaptation (LoRA). These models provide competitive performance with a fraction of the pixels and computational cost of full resolution non-foveated baselines, opening pathways for efficient and scalable active sensing for high-resolution egocentric vision. Code (https://github.com/nblauch/fovi) and pre-trained models (https://huggingface.co/fovi-pytorch) are available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf658cf0-5245-4572-8962-dcc28904fde8Builds on5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Early Convolutions Help Transformers See BetterTete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell et al.NeurIPS 2021 · 974 citations
- Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive GazingBaifeng Shi, Stephanie Fu, Long Lian, Hanrong Ye et al.CVPR 2026 · 9 citations
- Scaling Vision Pre-Training to 4K ResolutionBaifeng Shi, Boyi Li, Han Cai, Yao Lu et al.CVPR 2025
- FFCV: Accelerating Training by Removing Data BottlenecksGuillaume Leclerc, Andrew Ilyas, Logan Engstrom, Sung Min Park et al.CVPR 2023
Related papers
- Modeling Rapid Contextual Learning in the Visual Cortex with Fast-Weight Deep Autoencoder NetworksYue Li, Weifan Wang, Tai Sing LeeAAAI 2026
- Seeing More with Less: Human-like Representations in Vision ModelsAndrey Gizdov, Shimon Ullman, Daniel HarariCVPR 2025
- Dr. RAW: Towards General High-Level Vision from RAW with Efficient Task ConditioningWenjun Huang, Ziteng Cui, Yinqiang Zheng, Yirui He et al.NeurIPS 2025 · 5 citations
- FOVEA: Foveated Image Magnification for Autonomous NavigationChittesh Thavamani, Mengtian Li, Nicolas Cebron, Deva RamananICCV 2021 · 45 citations
- LiDeRe: A Lightweight Readout for Fast and Data-Efficient Dense PredictionTimo Lüddecke, Jan F. Meier, Jan van Delden, Alexander S. EckerCVPR 2026
