Gemel: Model Merging for Memory-Efficient, Real-Time Video Analytics at the Edge
Arthi Padmanabhan, Neil Agarwal, Anand P. Iyer, Ganesh Ananthanarayanan, Yuanchao Shu, Nikolaos Karianakis, Guoqing Harry Xu, Ravi Netravali
摘要
Video analytics pipelines have steadily shifted to edge deployments to reduce bandwidth overheads and privacy violations, but in doing so, face an ever-growing resource tension. Most notably, edge-box GPUs lack the memory needed to concurrently house the growing number of (increasingly complex) models for real-time inference. Unfortunately, existing solutions that rely on time/space sharing of GPU resources are insufficient as the required swapping delays result in unacceptable frame drops and accuracy violations. We present model merging, a new memory management technique that exploits architectural similarities between edge vision models by judiciously sharing their layers (including weights) to reduce workload memory costs and swapping delays. Our system, GEMEL, efficiently integrates merging into existing pipelines by (1) leveraging several guiding observations about per-model memory usage and inter-layer dependencies to quickly identify fruitful and accuracy-preserving merging configurations, and (2) altering edge inference schedules to maximize merging benefits. Experiments across diverse workloads reveal that GEMEL reduces memory usage by up to 60.7%, and improves overall accuracy by 8-39% relative to time/space sharing alone.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Known Knowns and Unknowns: Near-realtime Earth Observation Via Query Bifurcation in ServalBill Tao, Om Chabra, Ishani Janveja, Indranil Gupta 等NSDI 2024 · 被引用 44 次
- USHER: Holistic Interference Avoidance for Resource Optimized ML InferenceSudipta Saha Shubha, Haiying Shen, Anand P. IyerOSDI 2024 · 被引用 35 次
- Vulcan: Automatic Query Planning for Live ML AnalyticsYiwen Zhang, Xumiao Zhang, Ganesh Ananthanarayanan, Anand P. Iyer 等NSDI 2024 · 被引用 17 次
- Region-based Content Enhancement for Efficient Video Analytics at the EdgeWeijun Wang, Liang Mi, Shaowei Cen, Haipeng Dai 等NSDI 2025 · 被引用 12 次
- Gecko: Resource-Efficient and Accurate Queries in Real-Time Video Streams at the EdgeLiang Wang, Xiaoyang Qu, Jianzong Wang, Guokuan Li 等INFOCOM 2024 · 被引用 11 次
它引用的顶会 Paper15
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao 等OSDI 2020 · 被引用 392 次
- AdaShare: Learning What To Share For Efficient Deep Multi-Task LearningXimeng Sun, Rameswar Panda, Rogério Feris, Kate SaenkoNeurIPS 2020 · 被引用 337 次
- AntMan: Dynamic Scaling on GPU Clusters for Deep LearningWencong Xiao, Shiru Ren, Yong Li, Yang Zhang 等OSDI 2020 · 被引用 260 次
- ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learningSamyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith 等SC 2021 · 被引用 254 次
- Server-Driven Video Streaming for Deep Learning InferenceKuntai Du, Ahsan Pervaiz, Xin Yuan, Aakanksha Chowdhery 等SIGCOMM 2020 · 被引用 238 次
相关 Paper
- Remembrall: Leaning into Memory for Accurate Video Analytics on System-on-Chip GPUsMurali Ramanujam, Yinwei Dai, Kyle Jamieson, Ravi NetravaliNSDI 2026
- ResMap: Exploiting Sparse Residual Feature Map for Accelerating Cross-Edge Video AnalyticsNing Chen, Shuai Zhang, Sheng Zhang, Yuting Yan 等INFOCOM 2023 · 被引用 11 次
- S andhi : Fine-Grained Merging for Memory Efficient Multi-Model ServingVima Gupta, Oytun Kuday Duran, Nandan Suresh Meda, Ikhyun An 等SOSP 2026
- Cross-Camera Inference on the Constrained EdgeJingzong Li, Libin Liu, Hong Xu, Shudeng Wu 等INFOCOM 2023 · 被引用 31 次
- RECL: Responsive Resource-Efficient Continuous Learning for Video AnalyticsMehrdad Khani Shirkoohi, Ganesh Ananthanarayanan, Kevin Hsieh, Junchen Jiang 等NSDI 2023
