Cost-Aware Fine-Grained Recognition for IoTs Based on Sequential Fixations
Hanxiao Wang, Venkatesh Saligrama, Stan Sclaroff, Vitaly Ablavsky
摘要
We consider the problem of fine-grained classification on an edge camera device that has limited power. The edge device must sparingly interact with the cloud to minimize communication bits to conserve power, and the cloud upon receiving the edge inputs returns a classification label. To deal with fine-grained classification, we adopt the perspective of sequential fixation with a foveated field-of-view to model cloud-edge interactions. We propose a novel deep reinforcement learning-based foveation model, DRIFT, that sequentially generates and recognizes mixed-acuity images. Training of DRIFT requires only image-level category labels and encourages fixations to contain task-relevant information, while maintaining data efficiency. Specifically, we train a foveation actor network with a novel Deep Deterministic Policy Gradient by Conditioned Critic and Coaching (DDPGC3) algorithm. In addition, we propose to shape the reward to provide informative feedback after each fixation to better guide RL training. We demonstrate the effectiveness of DRIFT on this task by evaluating on five fine-grained classification benchmark datasets, and show that the proposed approach achieves state-of-the-art performance with over 3X reduction in transmitted pixels.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Debiasing Model Updates for Improving Personalized Federated TrainingDurmus Alp Emre Acar, Yue Zhao, Ruizhao Zhu, Ramon Matas Navarro 等ICML 2021 · 被引用 75 次
- SaccadeCam: Adaptive Visual Attention for Monocular Depth SensingBrevin Tilmon, Sanjeev J. KoppalICCV 2021 · 被引用 6 次
相关 Paper
- SteerCam: Multi-Camera Edge Perception via Dynamic Joint Steering & CollaborationDhanuja Wanniarachchige, Kasthuri Jayarajah, Dulaj Weerakoon, Tarek F. Abdelzaher 等INFOCOM 2026
- Visual hyperacuity with moving sensor and recurrent neural computationsAlexander Rivkind, Or Ram, Eldad Assa, Michael Kreiserman 等ICLR 2022 · 被引用 8 次
- FovRL: Joint Foveation and Quality Control for Immersive VR Streaming Using Reinforcement LearningYuk Hang Tsui, Ze Wu, Ahmad Alhilal, Matti Siekkinen 等WWW 2026 · 被引用 1 次
- Stabilizing and Accelerating Autofocus with Expert Trajectory Regularized Deep Reinforcement LearningShouhang Zhu, Chenglin Li, Yuankun Jiang, Li Wei 等CVPR 2025
- Coarse-to-Fine Q-attention: Efficient Learning for Visual Robotic Manipulation via DiscretisationStephen James, Kentaro Wada, Tristan Laidlow, Andrew J. DavisonCVPR 2022 · 被引用 65 次
