Allo: A Programming Model for Composable Accelerator Design
Hongzheng Chen, Niansong Zhang, Shaojie Xiang, Zhichen Zeng, Mengjia Dai, Zhiru Zhang
摘要
Special-purpose hardware accelerators are increasingly pivotal for sustaining performance improvements in emerging applications, especially as the benefits of technology scaling continue to diminish. However, designers currently lack effective tools and methodologies to construct complex, high-performance accelerator architectures in a productive manner. Existing high-level synthesis (HLS) tools often require intrusive sourcelevel changes to attain satisfactory quality of results. Despite the introduction of several new accelerator design languages (ADLs) aiming to enhance or replace HLS, their advantages are more evident in relatively simple applications with a single kernel. Existing ADLs prove less effective for realistic hierarchical designs with multiple kernels, even if the design hierarchy is flattened.
In this paper, we introduce Allo, a composable programming model for efficient spatial accelerator design. Allo decouples hardware customizations, including compute, memory, communication, and data type from algorithm specification, and encapsulates them as a set of customization primitives. Allo preserves the hierarchical structure of an input program by combining customizations from different functions in a bottomup, type-safe manner. This approach facilitates holistic optimizations that span across function boundaries. We conduct comprehensive experiments on commonly-used HLS benchmarks and several realistic deep learning models. Our evaluation shows that Allo can outperform state-of-the-art HLS tools and ADLs on all test cases in the PolyBench. For the GPT2 model, the inference latency of the Allo generated accelerator is 1.7× faster than the NVIDIA A100 GPU with 5.4× higher energy efficiency, demonstrating the capability of Allo to handle large-scale designs. * Equal contribution. † Work was done when Zhichen and Mengjia interned at Cornell.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- HeuriGym: An Agentic Benchmark for LLM-Crafted Heuristics in Combinatorial OptimizationHongzheng Chen, Yingheng Wang, Yaohui Cai, Hins Hu 等ICLR 2026 · 被引用 26 次
- SmoothE: Differentiable E-Graph ExtractionYaohui Cai, Kaixin Yang, Chenhui Deng, Cunxi Yu 等ASPLOS 2025 · 被引用 12 次
- StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMsHanchen Ye, Deming ChenMICRO 2025 · 被引用 5 次
- ChatHLS: Towards Systematic Design Automation and Optimization for High-Level SynthesisRunkai Li, Jia Xiong, Xiuyuan He, Jieru Zhao 等ACL 2026 · 被引用 3 次
- OmniSim: Simulating Hardware with C Speed and RTL Accuracy for High-Level Synthesis DesignsRishov Sarkar, Cong HaoMICRO 2025 · 被引用 3 次
它引用的顶会 Paper17
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein 等ASPLOS 2024 · 被引用 693 次
- Ansor: Generating High-Performance Tensor Programs for Deep LearningLianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu 等OSDI 2020 · 被引用 551 次
- Gemmini: Enabling Systematic Deep-Learning Architecture Evaluation via Full-Stack IntegrationHasan Genc, Seah Kim, Alon Amid, Ameer Haj-Ali 等DAC 2021 · 被引用 325 次
- DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text GenerationSeongmin Hong, Seungjae Moon, Junsoo Kim, Sungjae Lee 等MICRO 2022 · 被引用 107 次
相关 Paper
- DSAGEN: Synthesizing Programmable Spatial AcceleratorsJian Weng, Sihao Liu, Vidushi Dadu, Zhengrong Wang 等ISCA 2020 · 被引用 140 次
- OverGen: Improving FPGA Usability through Domain-specific Overlay GenerationSihao Liu, Jian Weng, Dylan Kupsh, Atefeh Sohrabizadeh 等MICRO 2022 · 被引用 32 次
- ScaleHLS: A New Scalable High-Level Synthesis Framework on Multi-Level Intermediate RepresentationHanchen Ye, Cong Hao, Jianyi Cheng, Hyunmin Jeong 等HPCA 2022 · 被引用 77 次
- AMOS: enabling automatic mapping for tensor computations on spatial accelerators with hardware abstractionSize Zheng, Renze Chen, Anjiang Wei, Yicheng Jin 等ISCA 2022 · 被引用 63 次
- Predictable accelerator design with time-sensitive affine typesRachit Nigam, Sachille Atapattu, Samuel Thomas, Zhijing Li 等PLDI 2020 · 被引用 58 次
