Cost-Aware Fine-Grained Recognition for IoTs Based on Sequential Fixations
Hanxiao Wang, Venkatesh Saligrama, Stan Sclaroff, Vitaly Ablavsky
Abstract
We consider the problem of fine-grained classification on an edge camera device that has limited power. The edge device must sparingly interact with the cloud to minimize communication bits to conserve power, and the cloud upon receiving the edge inputs returns a classification label. To deal with fine-grained classification, we adopt the perspective of sequential fixation with a foveated field-of-view to model cloud-edge interactions. We propose a novel deep reinforcement learning-based foveation model, DRIFT, that sequentially generates and recognizes mixed-acuity images. Training of DRIFT requires only image-level category labels and encourages fixations to contain task-relevant information, while maintaining data efficiency. Specifically, we train a foveation actor network with a novel Deep Deterministic Policy Gradient by Conditioned Critic and Coaching (DDPGC3) algorithm. In addition, we propose to shape the reward to provide informative feedback after each fixation to better guide RL training. We demonstrate the effectiveness of DRIFT on this task by evaluating on five fine-grained classification benchmark datasets, and show that the proposed approach achieves state-of-the-art performance with over 3X reduction in transmitted pixels.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Debiasing Model Updates for Improving Personalized Federated TrainingDurmus Alp Emre Acar, Yue Zhao, Ruizhao Zhu, Ramon Matas Navarro et al.ICML 2021 · 75 citations
- SaccadeCam: Adaptive Visual Attention for Monocular Depth SensingBrevin Tilmon, Sanjeev J. KoppalICCV 2021 · 6 citations
Related papers
- SteerCam: Multi-Camera Edge Perception via Dynamic Joint Steering & CollaborationDhanuja Wanniarachchige, Kasthuri Jayarajah, Dulaj Weerakoon, Tarek F. Abdelzaher et al.INFOCOM 2026
- Visual hyperacuity with moving sensor and recurrent neural computationsAlexander Rivkind, Or Ram, Eldad Assa, Michael Kreiserman et al.ICLR 2022 · 8 citations
- FovRL: Joint Foveation and Quality Control for Immersive VR Streaming Using Reinforcement LearningYuk Hang Tsui, Ze Wu, Ahmad Alhilal, Matti Siekkinen et al.WWW 2026 · 1 citation
- Stabilizing and Accelerating Autofocus with Expert Trajectory Regularized Deep Reinforcement LearningShouhang Zhu, Chenglin Li, Yuankun Jiang, Li Wei et al.CVPR 2025
- Coarse-to-Fine Q-attention: Efficient Learning for Visual Robotic Manipulation via DiscretisationStephen James, Kentaro Wada, Tristan Laidlow, Andrew J. DavisonCVPR 2022 · 65 citations
