Policy-based Foveated Imaging and Perception
Howard Xiao, Jan Ackermann, Boyang Deng, Gordon Wetzstein
Abstract
Ultra-high-resolution image sensors offer the potential to capture fine spatial details critical for many visual perception tasks, but acquiring and processing all pixels at full resolution is often infeasible under realistic bandwidth, latency, and power constraints. Existing approaches address this challenge through acquisition strategies such as spatial or temporal downsampling, which irrevocably discard information before task relevance can be assessed. In this work, we introduce a real-time, predictive, and task-aware foveated imaging system that operates directly at image acquisition time. Leveraging emerging dual-stream sensor architectures, our method dynamically allocates limited pixel bandwidth to task-relevant regions of interest while maintaining a low-resolution global context. We formulate foveated acquisition as a sensor attention policy–learning problem, in which past observations guide actions that determine future measurements, closing the perception–acquisition loop. Through extensive simulation across multiple perception tasks, we demonstrate that our approach achieves high task performance under strict pixel budgets and significantly outperforms relevant baselines operating at the same bandwidth. We further validate our system on a 200-megapixel dual-stream sensor, capturing real-world videos under realistic bandwidth and latency constraints, demonstrating the practical feasibility of task-driven, acquisition-time foveated imaging. Our project website is at https://howardxiao.ca/foveated/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b0eb0e2c-1700-4af5-b75b-2b7b49162c86Builds on8
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- MVSNeRF: Fast Generalizable Radiance Field Reconstruction from Multi-View StereoAnpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang et al.ICCV 2021 · 1,024 citations
- Learning To Explore Using Active Neural SLAMDevendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta et al.ICLR 2020 · 603 citations
- MixFormerV2: Efficient Fully Transformer TrackingYutao Cui, Tianhui Song, Gangshan Wu, Limin WangNeurIPS 2023 · 193 citations
- Towards Attention-aware Foveated RenderingBrooke Krajancich, Petr Kellnhofer, Gordon WetzsteinSIGGRAPH 2023 · 41 citations
Related papers
- FOVEA: Foveated Image Magnification for Autonomous NavigationChittesh Thavamani, Mengtian Li, Nicolas Cebron, Deva RamananICCV 2021 · 45 citations
- The Perceptual Bandwidth Bottleneck in Vision-Language Models: Active Visual Reasoning via Sequential Experimental DesignAnjie Liu, Ziqin Gong, Yan Song, Yuxiang Chen et al.ICML 2026 · 1 citation
- FOCAS: Practical Video Super Resolution using Foveated RenderingLingdong Wang, Mohammad H. Hajiesmaili, Ramesh K. SitaramanACM MM 2021 · 21 citations
- Seeing More with Less: Human-like Representations in Vision ModelsAndrey Gizdov, Shimon Ullman, Daniel HarariCVPR 2025
- Saliency-Guided Foveated Video Encoding for Low-Latency and Immersive Cloud VRZe Wu, Ahmad Alhilal, Yuk Hang Tsui, Wen Jye Chai et al.IEEE VR 2026
