CAMEL: Co-Designing AI Models and eDRAMs for Efficient On-Device Learning
Sai Qian Zhang, Thierry Tambe, Nestor Cuevas, Gu-Yeon Wei, David Brooks
Abstract
On-device learning allows AI models to adapt to user data, thereby enhancing service quality on edge platforms. However, training AI on resource-limited devices poses significant challenges due to the demanding computing workload and the substantial memory consumption and data access required by deep neural networks (DNNs). To address these issues, we propose utilizing embedded dynamic random-access memory (eDRAM) as the primary storage medium for transient training data. In comparison to static random-access memory (SRAM), eDRAM provides higher storage density and lower leakage power, resulting in reduced access cost and power leakage. Nevertheless, to maintain the integrity of the stored data, periodic power-hungry refresh operations could potentially degrade system performance.
To minimize the occurrence of expensive eDRAM refresh operations, it is beneficial to shorten the lifetime of stored data during the training process. To achieve this, we adopt the principles of algorithm and hardware co-design, introducing a family of reversible DNN architectures that effectively decrease data lifetime and storage costs throughout training. Additionally, we present a highly efficient on-device training engine named CAMEL, which leverages eDRAM as the primary on-chip memory. This engine enables efficient on-device training with significantly reduced memory usage and off-chip DRAM traffic while maintaining superior training accuracy. We evaluate our CAMEL system on multiple DNNs with different datasets, demonstrating a 2.5× speedup of the training process and 2.8× training energy savings than the other baseline hardware platforms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 356cbb78-9741-4ccb-87c1-4aa971aa4848Cited by top-tier papers4
- PICACHU: Plug-In CGRA Handling Upcoming Nonlinear Operations in LLMsJiajun Qin, Tianhua Xia, Cheng Tan, Jeff Zhang et al.ASPLOS 2025 · 17 citations
- DACAPO: Accelerating Continuous Learning in Autonomous Systems for Video AnalyticsYoonsung Kim, Changhun Oh, Jinwoo Hwang, Wonung Kim et al.ISCA 2024 · 13 citations
- Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge ComputingTianhua Xia, Sai Qian ZhangMICRO 2025 · 2 citations
- RPU - A Reasoning Processing UnitMatthew Joseph Adiletta, Gu-Yeon Wei, David BrooksHPCA 2026 · 2 citations
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella et al.HPCA 2020 · 490 citations
- TensorDash: Exploiting Sparsity to Accelerate Deep Neural Network TrainingMostafa Mahmoud, Isak Edo, Ali Hadi Zadeh, Omar Mohamed Awad et al.MICRO 2020 · 78 citations
Related papers
- Efficient Memory Integration: MRAM-SRAM Hybrid Accelerator for Sparse On-Device LearningFan Zhang, Amitesh Sridharan, Wilman Tsai, Yiran Chen et al.DAC 2024 · 7 citations
- HH-PIM: Dynamic Optimization of Power and Performance with Heterogeneous-Hybrid PIM for Edge AI DevicesSangmin Jeon, Kangju Lee, Kyeongwon Lee, Woojoo LeeDAC 2025 · 4 citations
- CREAM: computing in ReRAM-assisted energy and area-efficient SRAM for neural network accelerationLiukai Xu, Songyuan Liu, Zhi Li, Dengfeng Wang et al.DAC 2022 · 6 citations
- Efficient On-Device Training via Gradient FilteringYuedong Yang, Guihong Li, Radu MarculescuCVPR 2023
- On-Device Training Under 256KB MemoryJi Lin, Ligeng Zhu, Wei-Ming Chen, Wei-Chen Wang et al.NeurIPS 2022 · 345 citations
