Towards Exact Gradient-based Training on Analog In-memory Computing
Zhaoxian Wu, Tayfun Gokmen, Malte J. Rasch, Tianyi Chen
Abstract
Given the high economic and environmental costs of using large vision or language models, analog in-memory accelerators present a promising solution for energy-efficient AI. While inference on analog accelerators has been studied recently, the training perspective is underexplored. Recent studies have shown that the"workhorse"of digital AI training - stochastic gradient descent (SGD) algorithm converges inexactly when applied to model training on non-ideal devices. This paper puts forth a theoretical foundation for gradient-based training on analog devices. We begin by characterizing the non-convergent issue of SGD, which is caused by the asymmetric updates on the analog devices. We then provide a lower bound of the asymptotic error to show that there is a fundamental performance limit of SGD-based analog training rather than an artifact of our analysis. To address this issue, we study a heuristic analog algorithm called Tiki-Taka that has recently exhibited superior empirical performance compared to SGD and rigorously show its ability to exactly converge to a critical point and hence eliminates the asymptotic error. The simulations verify the correctness of the analyses.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 601c554d-5d9a-420b-a124-e7e2d9e798b6Cited by top-tier papers2
- Analog In-memory Training on General Non-ideal Resistive Elements: The Impact of Response FunctionsZhaoxian Wu, Quan Xiao, Tayfun Gokmen, Omobayode Fagbohungbe et al.NeurIPS 2025 · 8 citations
- Dynamic Symmetric Point Tracking: Tackling Non-ideal Reference in Analog In-memory TrainingQuan Xiao, jindan li, Zhaoxian Wu, Tayfun Gokmen et al.ICML 2026
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Improved Analysis of Stochastic Gradient Descent with MomentumYanli Liu, Yuan Gao, Wotao YinNeurIPS 2020 · 328 citations
- A Guide Through the Zoo of Biased SGDYury Demidovich, Grigory Malinovsky, Igor Sokolov, Peter RichtárikNeurIPS 2023 · 56 citations
- Energy-based learning algorithms for analog computing: a comparative studyBenjamin Scellier, Maxence Ernoult, Jack D. Kendall, Suhas KumarNeurIPS 2023 · 54 citations
- Three-Way Trade-Off in Multi-Objective Learning: Optimization, Generalization and Conflict-AvoidanceLisha Chen, Heshan Devaka Fernando, Yiming Ying, Tianyi ChenNeurIPS 2023 · 53 citations
Related papers
- Towards training digitally-tied analog blocks via hybrid gradient computationTimothy Nest, Maxence ErnoultNeurIPS 2024 · 6 citations
- Analog Foundation ModelsJulian Büchel, Iason Chalas, Giovanni Acampa, An Chen et al.NeurIPS 2025 · 7 citations
- Accelerated On-Device Forward Neural Network Training with Module-Wise Descending AsynchronismXiaohan Zhao, Hualin Zhang, Zhouyuan Huo, Bin GuNeurIPS 2023 · 1 citation
- A fast algorithm to simulate nonlinear resistive networksBenjamin ScellierICML 2024 · 8 citations
- ePC: Fast and Deep Predictive Coding in Digital SimulationCédric Goemaere, Gaspard Oliviers, Rafal Bogacz, Thomas DemeesterICML 2026 · 3 citations
