FOVEA: Foveated Image Magnification for Autonomous Navigation
Chittesh Thavamani, Mengtian Li, Nicolas Cebron, Deva Ramanan
Abstract
Efficient processing of high-res video streams is safetycritical for many robotics applications such as autonomous driving. To maintain real-time performance, many practical systems downsample the video stream. But this can hurt downstream tasks such as (small) object detection. Instead, we take inspiration from biological vision systems that allocate more foveal "pixels" to salient parts of the scene. We introduce FOVEA, an approach for intelligent downsampling that ensures salient image regions remain "magnified" in the downsampled output. Given a high-res image, FOVEA applies a differentiable resampling layer that outputs a small fixed-size image canvas, which is then processed with a differentiable vision module (e.g., object detection network), whose output is then differentiably backward mapped onto the original image size. The key idea is to resample such that background pixels can make room for salient pixels of interest. In order to ensure the overall pipeline remains efficient, FOVEA makes use of cheap and readily available cues for saliency, including datasetspecific spatial priors or temporal priors computed from object predictions in the recent past. On the autonomous driving datasets Argoverse-HD and BDD100K, our proposed method boosts the detection AP over standard Faster R-CNN, both with and without finetuning. Without any noticeable increase in compute, we improve accuracy on small objects by over 2x without degrading performance on large objects. Finally, FOVEA sets a new record for streaming AP (from 17.8 to 23.0 on a GTX 1080 Ti GPU), a metric designed to capture both accuracy and latency. * denotes equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c15580ed-7abe-4abd-b16a-a53d167c2442Cited by top-tier papers9
- ZoomTrack: Target-aware Non-uniform Resizing for Efficient Visual TrackingYutong Kou, Jin Gao, Bing Li, Gang Wang et al.NeurIPS 2023 · 74 citations
- Real-time Object Detection for Streaming PerceptionJinrong Yang, Songtao Liu, Zeming Li, Xiaoping Li et al.CVPR 2022 · 61 citations
- A Dual-Stream Neural Network Explains the Functional Segregation of Dorsal and Ventral Visual Pathways in Human BrainsMinkyu Choi, Kuan Han, Xiaokai Wang, Yizhen Zhang et al.NeurIPS 2023 · 33 citations
- A Partially-Supervised Reinforcement Learning Framework for Visual Active SearchAnindya Sarkar, Nathan Jacobs, Yevgeniy VorobeychikNeurIPS 2023 · 13 citations
- Chanakya: Learning Runtime Decisions for Adaptive Real-Time PerceptionAnurag Ghosh, Vaibhav Balloli, Akshay Nambi, Aditya Singh et al.NeurIPS 2023 · 10 citations
Builds on11
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 1,030 citations
- Efficient Segmentation: Learning Downsampling Near Semantic BoundariesDmitrii Marin, Zijian He, Peter Vajda, Priyam Chatterjee et al.ICCV 2019 · 107 citations
- Budgeted Training: Rethinking Deep Neural Network Training Under Resource ConstraintsMengtian Li, Ersin Yumer, Deva RamananICLR 2020 · 58 citations
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora et al.CVPR 2020
Related papers
- Learned Two-Plane Perspective Prior based Image Resampling for Efficient Object DetectionAnurag Ghosh, N. Dinesh Reddy, Christoph Mertz, Srinivasa G. NarasimhanCVPR 2023
- FOCAS: Practical Video Super Resolution using Foveated RenderingLingdong Wang, Mohammad H. Hajiesmaili, Ramesh K. SitaramanACM MM 2021 · 21 citations
- Policy-based Foveated Imaging and PerceptionHoward Xiao, Jan Ackermann, Boyang Deng, Gordon WetzsteinSIGGRAPH 2026
- AdaSpot: Spend Resolution Where It Matters for Precise Event SpottingArtur Xarles, Sergio Escalera, Thomas B. Moeslund, Albert ClapésCVPR 2026 · 2 citations
- SpotStream: Real-Time Video Transmission for Autonomous Driving via Small Object-Aware ROIZelin Song, Huanhuan Zhang, Pengcheng Zhang, Mingyue Zhao et al.INFOCOM 2026
