Calculon: a methodology and tool for high-level co-design of systems and large language models
Mikhail Isaev, Nic McDonald, Larry Dennison, Richard W. Vuduc
摘要
This paper presents a parameterized analytical performance model of transformer-based Large Language Models (LLMs) for guiding high-level algorithm-architecture codesign studies. This model derives from an extensive survey of performance optimizations that have been proposed for the training and inference of LLMs; the model's parameters capture application characteristics, the hardware system, and the space of implementation strategies. With such a model, we can systematically explore a joint space of hardware and software configurations to identify optimal system designs under given constraints, like the total amount of system memory. We implemented this model and methodology in a Python-based open-source tool called Calculon. Using it, we identified novel system designs that look significantly different from current inference and training systems, showing quantitatively the estimated potential to achieve higher efficiency, lower cost, and better scalability.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper8
- Forecasting GPU Performance for Deep Learning Training and InferenceSeonho Lee, Amar Phanishayee, Divya MahajanASPLOS 2025 · 被引用 31 次
- vTrain: A Simulation Framework for Evaluating Cost-Effective and Compute-Optimal Large Language Model TrainingJehyeon Bang, Yujeong Choi, Myeongwoo Kim, Yongdeok Kim 等MICRO 2024 · 被引用 16 次
- FlowCheck: Decoupling Checkpointing and Training of Large-Scale ModelsZimeng Huang, Hao Nie, Haonan Jia, Bo Jiang 等EuroSys 2025 · 被引用 3 次
- Scalable Synthesis of Distributed Llm Workloads Through Symbolic Tensor GraphsChanghai Man, Joongun Park, Hanjiang Wu, Huan Xu 等ISCA 2026 · 被引用 2 次
- WATOS: Efficient LLM Training Strategies and Architecture Co-Exploration for Wafer-Scale ChipHuizheng Wang, Zichuan Wang, Hongbin Wang, Jingxiang Hou 等HPCA 2026 · 被引用 2 次
相关 Paper
- LLMShare: Optimizing LLM Inference Serving with Hardware Architecture ExplorationHongduo Liu, Chen Bai, Peng Xu, Lihao Yin 等DAC 2025 · 被引用 2 次
- Cheaply Estimating Inference Efficiency Metrics for Autoregressive Transformer ModelsDeepak Narayanan, Keshav Santhanam, Peter Henderson, Rishi Bommasani 等NeurIPS 2023 · 被引用 14 次
- LLM-Pilot: Characterize and Optimize Performance of your LLM Inference ServicesMalgorzata Lazuka, Andreea Anghel, Thomas P. ParnellSC 2024 · 被引用 17 次
- Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMsSong Bian, Tao Yu, Shivaram Venkataraman, Youngsuk ParkICLR 2026 · 被引用 3 次
- Optimas: An Intelligent Analytics-Informed Generative AI Framework for Performance OptimizationMohammad Zaeed, Tanzima Z. Islam, Vladimir IndicKDD 2026
