DS-LLM: Leveraging Dynamical Systems to Enhance Both Training and Inference of Large Language Models
Ruibing Song, Chuan Liu, Chunshu Wu, Ang Li, Dongfang Liu, Ying Nian Wu, Tong Geng
摘要
The training of large language models (LLMs) faces significant computational cost challenges, limiting their scalability toward artificial general intelligence (AGI) and broader adoption. With model sizes doubling approximately every 3.4 months and training costs escalating from 64 million USD for GPT-4 in 2020 to 191 million USD for Gemini Ultra in 2023, the economic burden has become unsustainable. While techniques such as quantization offer incremental improvements, they fail to address the fundamental computational bottleneck. In this work, we introduce DS-LLM, a novel framework that leverages dynamical system (DS)-based machines, which exploit Natural Annealing to rapidly converge to minimal energy states, yielding substantial efficiency gains. Unlike traditional methods, DS-LLM maps LLM components to optimization problems solvable via Hamiltonian configurations and utilizes continuous electric current flow in DS-machines for hardware-native gradient descent during training. We mathematically demonstrate the equivalence between conventional LLMs and DS-LLMs and present a method for transforming a trained LLM into a DS-LLM. Experimental evaluations across multiple model sizes demonstrate orders-of-magnitude improvements in speed and energy efficiency for both training and inference while maintaining consistent accuracy. Additionally, we provide an in-depth analysis of the challenges and potential solutions associated with this emerging computing paradigm, aiming to lay a solid foundation for future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric TasksWenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu 等NeurIPS 2023 · 被引用 725 次
- BRIM: Bistable Resistively-Coupled Ising MachineRichard Afoakwa, Yiqiao Zhang, Uday Kumar Reddy Vengalam, Zeljko Ignjatovic 等HPCA 2021 · 被引用 57 次
- Cambricon-LLM: A Chiplet-Based Hybrid Architecture for On-Device Inference of 70B LLMZhongkai Yu, Shengwen Liang, Tianyun Ma, Yunke Cai 等MICRO 2024 · 被引用 29 次
- Increasing ising machine capacity with multi-chip architecturesAnshujit Sharma, Richard Afoakwa, Zeljko Ignjatovic, Michael C. HuangISCA 2022 · 被引用 28 次
相关 Paper
- Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLMTianyi Wu, Xiaoxi Sun, Yanhua Jiao, Yulin Li 等ICLR 2026 · 被引用 6 次
- DS-TPU: Dynamical System for on-Device Lifelong Graph Learning with Nonlinear Node InteractionChunshu Wu, Ruibing Song, Chuan Liu, Pouya Haghi 等ISCA 2025 · 被引用 3 次
- Beyond Gemini-3-Pro: Revisiting LLM Routing and Aggregation at ScaleShengji Tang, Weihao Lin, Peng Ye, Jingqi Ye 等ICML 2026
- InstaTrain: Adaptive Training via Ultra-Fast Natural Annealing within Dynamical SystemsChuan Liu, Ruibing Song, Chunshu Wu, Pouya Haghi 等ICLR 2025
- BitDP: Ultra-low-bit Communication for Data Parallelism in LLM TrainingXiaozhe Ren, Qiong LuoAAAI 2026
