Improved Efficiency Based on Learned Saccade and Continuous Scene Reconstruction From Foveated Visual Sampling
Jiayang Liu, Yiming Bu, Daniel Tso, Qinru Qiu
Abstract
High accuracy, low latency and high energy efficiency represent a set of conflicting goals when searching for system solutions for image classification and detection. While high-quality images naturally result in more precise detection and classification, they also result in a heavier computational workload for imaging and processing, reduced camera frame rates, and increased data communication between the camera and processor. Taking inspiration from the foveal-peripheral sampling mechanism, and saccade mechanism of the human visual system and the filling-in phenomena of brain, we have developed an active scene reconstruction architecture based on multiple foveal views. This model stitches together information from a sequence of foveal-peripheral views, which are sampled from multiple glances. Assisted by a reinforcement learning-based saccade mechanism, our model reduces the required input pixels by over 90% per frame while maintaining the same level of performance in image recognition as with the original images. We evaluated the effectiveness of our model using the GTSRB dataset and the ImageNet dataset. Using an equal number of input pixels, our model demonstrates a 5% higher image recognition accuracy compared to state-of-theart foveal-peripheral based vision systems. Furthermore, we demonstrate that our foveal sampling/saccadic scene reconstruction model exhibits significantly lower complexity and higher data efficiency during the training phase compared to existing approaches. Code is available at Github.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Biologically Inspired Mechanisms for Adversarial RobustnessManish V. Reddy, Andrzej Banburski, Nishka Pant, Tomaso A. PoggioNeurIPS 2020 · 53 citations
- Finding Biological Plausibility for Adversarially Robust Features via Metameric TasksAnne Harrington, Arturo DezaICLR 2022 · 23 citations
- Learning When and Where to Zoom With Deep Reinforcement LearningBurak Uzkent, Stefano ErmonCVPR 2020
- Masked Autoencoders Are Scalable Vision LearnersKaiming He, Xinlei Chen, Saining Xie, Yanghao Li et al.CVPR 2022
Related papers
- FOCAS: Practical Video Super Resolution using Foveated RenderingLingdong Wang, Mohammad H. Hajiesmaili, Ramesh K. SitaramanACM MM 2021 · 21 citations
- Cost-Aware Fine-Grained Recognition for IoTs Based on Sequential FixationsHanxiao Wang, Venkatesh Saligrama, Stan Sclaroff, Vitaly AblavskyICCV 2019 · 2 citations
- High-Speed Image Reconstruction Through Short-Term Plasticity for Spiking CamerasYajing Zheng, Lingxiao Zheng, Zhaofei Yu, Boxin Shi et al.CVPR 2021
- SaccadeCam: Adaptive Visual Attention for Monocular Depth SensingBrevin Tilmon, Sanjeev J. KoppalICCV 2021 · 6 citations
- Recognizing High-Speed Moving Objects with Spike CameraJunwei Zhao, Jianming Ye, Shiliang Zhang, Zhaofei Yu et al.ACM MM 2023 · 4 citations
