Gemel: Model Merging for Memory-Efficient, Real-Time Video Analytics at the Edge
Arthi Padmanabhan, Neil Agarwal, Anand P. Iyer, Ganesh Ananthanarayanan, Yuanchao Shu, Nikolaos Karianakis, Guoqing Harry Xu, Ravi Netravali
Abstract
Video analytics pipelines have steadily shifted to edge deployments to reduce bandwidth overheads and privacy violations, but in doing so, face an ever-growing resource tension. Most notably, edge-box GPUs lack the memory needed to concurrently house the growing number of (increasingly complex) models for real-time inference. Unfortunately, existing solutions that rely on time/space sharing of GPU resources are insufficient as the required swapping delays result in unacceptable frame drops and accuracy violations. We present model merging, a new memory management technique that exploits architectural similarities between edge vision models by judiciously sharing their layers (including weights) to reduce workload memory costs and swapping delays. Our system, GEMEL, efficiently integrates merging into existing pipelines by (1) leveraging several guiding observations about per-model memory usage and inter-layer dependencies to quickly identify fruitful and accuracy-preserving merging configurations, and (2) altering edge inference schedules to maximize merging benefits. Experiments across diverse workloads reveal that GEMEL reduces memory usage by up to 60.7%, and improves overall accuracy by 8-39% relative to time/space sharing alone.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- Known Knowns and Unknowns: Near-realtime Earth Observation Via Query Bifurcation in ServalBill Tao, Om Chabra, Ishani Janveja, Indranil Gupta et al.NSDI 2024 · 44 citations
- USHER: Holistic Interference Avoidance for Resource Optimized ML InferenceSudipta Saha Shubha, Haiying Shen, Anand P. IyerOSDI 2024 · 35 citations
- Vulcan: Automatic Query Planning for Live ML AnalyticsYiwen Zhang, Xumiao Zhang, Ganesh Ananthanarayanan, Anand P. Iyer et al.NSDI 2024 · 17 citations
- Region-based Content Enhancement for Efficient Video Analytics at the EdgeWeijun Wang, Liang Mi, Shaowei Cen, Haipeng Dai et al.NSDI 2025 · 12 citations
- Gecko: Resource-Efficient and Accurate Queries in Real-Time Video Streams at the EdgeLiang Wang, Xiaoyang Qu, Jianzong Wang, Guokuan Li et al.INFOCOM 2024 · 11 citations
Builds on15
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao et al.OSDI 2020 · 392 citations
- AdaShare: Learning What To Share For Efficient Deep Multi-Task LearningXimeng Sun, Rameswar Panda, Rogério Feris, Kate SaenkoNeurIPS 2020 · 337 citations
- AntMan: Dynamic Scaling on GPU Clusters for Deep LearningWencong Xiao, Shiru Ren, Yong Li, Yang Zhang et al.OSDI 2020 · 260 citations
- ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learningSamyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith et al.SC 2021 · 254 citations
- Server-Driven Video Streaming for Deep Learning InferenceKuntai Du, Ahsan Pervaiz, Xin Yuan, Aakanksha Chowdhery et al.SIGCOMM 2020 · 238 citations
Related papers
- Remembrall: Leaning into Memory for Accurate Video Analytics on System-on-Chip GPUsMurali Ramanujam, Yinwei Dai, Kyle Jamieson, Ravi NetravaliNSDI 2026
- ResMap: Exploiting Sparse Residual Feature Map for Accelerating Cross-Edge Video AnalyticsNing Chen, Shuai Zhang, Sheng Zhang, Yuting Yan et al.INFOCOM 2023 · 11 citations
- S andhi : Fine-Grained Merging for Memory Efficient Multi-Model ServingVima Gupta, Oytun Kuday Duran, Nandan Suresh Meda, Ikhyun An et al.SOSP 2026
- Cross-Camera Inference on the Constrained EdgeJingzong Li, Libin Liu, Hong Xu, Shudeng Wu et al.INFOCOM 2023 · 31 citations
- RECL: Responsive Resource-Efficient Continuous Learning for Video AnalyticsMehrdad Khani Shirkoohi, Ganesh Ananthanarayanan, Kevin Hsieh, Junchen Jiang et al.NSDI 2023
