Remembrall: Leaning into Memory for Accurate Video Analytics on System-on-Chip GPUs
Murali Ramanujam, Yinwei Dai, Kyle Jamieson, Ravi Netravali
Abstract
Continually retraining models has emerged as a primary technique to enable high-accuracy video analytics on edge devices. Yet, existing systems employ such adaptation by relying on the spare compute resources that traditional (memoryconstrained) edge servers afford. In contrast, mobile edge devices such as drones and dashcams offer a fundamentally different resource profile: weak(er) compute with abundant unified memory pools. We present Remembrall, a continuous learning system for the mobile edge's System-on-Chip GPUs. Our driving insight is that visually distinct scenes that require retraining exhibit substantial overlap in model embeddings; if captured into a base model on device memory, specializing to each new scene can become lightweight, requiring very few samples. To practically realize this approach, Remembrall presents new, compute-efficient techniques to (1) select high-utility data samples for retraining specialized models, (2) update the base model without complete retraining, and (3) time-share compute resources between retraining and live inference for maximal accuracy. Across diverse workloads, Remembrall lowers retraining costs by 2.8-10× compared to existing systems, resulting in 18-45% higher accuracies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 006c0571-2090-4fa4-b558-1506cb1abab0Builds on14
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 994 citations
- Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAMLAniruddh Raghu, Maithra Raghu, Samy Bengio, Oriol VinyalsICLR 2020 · 736 citations
- SqueezeLLM: Dense-and-Sparse QuantizationSehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong et al.ICML 2024 · 306 citations
- AntMan: Dynamic Scaling on GPU Clusters for Deep LearningWencong Xiao, Shiru Ren, Yong Li, Yang Zhang et al.OSDI 2020 · 260 citations
Related papers
- RECL: Responsive Resource-Efficient Continuous Learning for Video AnalyticsMehrdad Khani Shirkoohi, Ganesh Ananthanarayanan, Kevin Hsieh, Junchen Jiang et al.NSDI 2023
- Ekya: Continuous Learning of Video Analytics Models on Edge Compute ServersRomil Bhardwaj, Zhengxu Xia, Ganesh Ananthanarayanan, Junchen Jiang et al.NSDI 2022
- SkyCL: Swift Continuous Learning with Kinship-Awareness for Multi-Drone Video Analytics under Drastic DriftYuanzheng Tan, Qing Li, Jiaqi Cui, Junkun Peng et al.WWW 2026
- DACAPO: Accelerating Continuous Learning in Autonomous Systems for Video AnalyticsYoonsung Kim, Changhun Oh, Jinwoo Hwang, Wonung Kim et al.ISCA 2024 · 13 citations
- Gemel: Model Merging for Memory-Efficient, Real-Time Video Analytics at the EdgeArthi Padmanabhan, Neil Agarwal, Anand P. Iyer, Ganesh Ananthanarayanan et al.NSDI 2023 · 94 citations
