FOVEA: Foveated Image Magnification for Autonomous Navigation
Chittesh Thavamani, Mengtian Li, Nicolas Cebron, Deva Ramanan
摘要
Efficient processing of high-res video streams is safetycritical for many robotics applications such as autonomous driving. To maintain real-time performance, many practical systems downsample the video stream. But this can hurt downstream tasks such as (small) object detection. Instead, we take inspiration from biological vision systems that allocate more foveal "pixels" to salient parts of the scene. We introduce FOVEA, an approach for intelligent downsampling that ensures salient image regions remain "magnified" in the downsampled output. Given a high-res image, FOVEA applies a differentiable resampling layer that outputs a small fixed-size image canvas, which is then processed with a differentiable vision module (e.g., object detection network), whose output is then differentiably backward mapped onto the original image size. The key idea is to resample such that background pixels can make room for salient pixels of interest. In order to ensure the overall pipeline remains efficient, FOVEA makes use of cheap and readily available cues for saliency, including datasetspecific spatial priors or temporal priors computed from object predictions in the recent past. On the autonomous driving datasets Argoverse-HD and BDD100K, our proposed method boosts the detection AP over standard Faster R-CNN, both with and without finetuning. Without any noticeable increase in compute, we improve accuracy on small objects by over 2x without degrading performance on large objects. Finally, FOVEA sets a new record for streaming AP (from 17.8 to 23.0 on a GTX 1080 Ti GPU), a metric designed to capture both accuracy and latency. * denotes equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- ZoomTrack: Target-aware Non-uniform Resizing for Efficient Visual TrackingYutong Kou, Jin Gao, Bing Li, Gang Wang 等NeurIPS 2023 · 被引用 74 次
- Real-time Object Detection for Streaming PerceptionJinrong Yang, Songtao Liu, Zeming Li, Xiaoping Li 等CVPR 2022 · 被引用 61 次
- A Dual-Stream Neural Network Explains the Functional Segregation of Dorsal and Ventral Visual Pathways in Human BrainsMinkyu Choi, Kuan Han, Xiaokai Wang, Yizhen Zhang 等NeurIPS 2023 · 被引用 33 次
- A Partially-Supervised Reinforcement Learning Framework for Visual Active SearchAnindya Sarkar, Nathan Jacobs, Yevgeniy VorobeychikNeurIPS 2023 · 被引用 13 次
- Chanakya: Learning Runtime Decisions for Adaptive Real-Time PerceptionAnurag Ghosh, Vaibhav Balloli, Akshay Nambi, Aditya Singh 等NeurIPS 2023 · 被引用 10 次
它引用的顶会 Paper11
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 被引用 1,030 次
- Efficient Segmentation: Learning Downsampling Near Semantic BoundariesDmitrii Marin, Zijian He, Peter Vajda, Priyam Chatterjee 等ICCV 2019 · 被引用 107 次
- Budgeted Training: Rethinking Deep Neural Network Training Under Resource ConstraintsMengtian Li, Ersin Yumer, Deva RamananICLR 2020 · 被引用 58 次
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora 等CVPR 2020
相关 Paper
- Learned Two-Plane Perspective Prior based Image Resampling for Efficient Object DetectionAnurag Ghosh, N. Dinesh Reddy, Christoph Mertz, Srinivasa G. NarasimhanCVPR 2023
- FOCAS: Practical Video Super Resolution using Foveated RenderingLingdong Wang, Mohammad H. Hajiesmaili, Ramesh K. SitaramanACM MM 2021 · 被引用 21 次
- Policy-based Foveated Imaging and PerceptionHoward Xiao, Jan Ackermann, Boyang Deng, Gordon WetzsteinSIGGRAPH 2026
- AdaSpot: Spend Resolution Where It Matters for Precise Event SpottingArtur Xarles, Sergio Escalera, Thomas B. Moeslund, Albert ClapésCVPR 2026 · 被引用 2 次
- SpotStream: Real-Time Video Transmission for Autonomous Driving via Small Object-Aware ROIZelin Song, Huanhuan Zhang, Pengcheng Zhang, Mingyue Zhao 等INFOCOM 2026
