Taming the Zoo: The Unified GraphIt Compiler Framework for Novel Architectures
Ajay Brahmakshatriya, Emily Furst, Victor A. Ying, Claire Hsu, Changwan Hong, Max Ruttenberg, Yunming Zhang, Dai Cheol Jung, Dustin Richmond, Michael B. Taylor, Julian Shun, Mark Oskin
摘要
We live in a new Cambrian Explosion of hardware devices. The end of conventional processor scaling has driven research and industry practice to explore a new generation of approaches. The old DNA of architecture design, including vectors, threads, shared or private memories, coherence or message passing, dataflow or von Neumann execution, are hybridized together in new and exciting ways. Each new architecture exposes a unique hardware-level API. Performance and energy efficiency are critically dependent on how well programs can use these APIs. One approach is to implement custom libraries for each new hardware architecture and application domain. A more scalable approach is to utilize a portable compiler infrastructure tailored to the application domain that makes it easy to generate efficient code for a diverse set of architectures with minimal porting effort.
We propose the Unified GraphIt Compiler framework (UGC), which does exactly this for graph applications. UGC achieves portability with reasonable effort by decoupling the architectureindependent algorithm from the architecture-specific schedules and backends. We introduce a new domain-specific intermediate representation, GraphIR, that is key to this decoupling. GraphIR encodes high-level algorithm and optimization information needed for hardware-specific code generation, making it easy to develop different backends (GraphVMs) for diverse architectures, including CPUs, GPUs, and next-generation hardware such as Swarm and the HammerBlade manycore. We also build scheduling language extensions that make it easy to expose optimization decisions like load balancing strategies, blocking for locality, and other data structure choices. We evaluate UGC on five algorithms and 10 input graphs on these 4 distinct architectures and show that UGC enables implementing optimizations that can provide up to 53× speedup over programmer-generated straightforward implementations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SpZip: Architectural Support for Effective Data Compression In Irregular ApplicationsYifan Yang, Joel S. Emer, Daniel SánchezISCA 2021 · 被引用 31 次
- Scalable, Programmable and Dense: The HammerBlade Open-Source RISC-V ManycoreDai Cheol Jung, Max Ruttenberg, Paul Gao, Scott Davidson 等ISCA 2024 · 被引用 11 次
它引用的顶会 Paper1
相关 Paper
- uGrapher: High-Performance Graph Operator Computation via Unified Abstraction for Graph Neural NetworksYangjie Zhou, Jingwen Leng, Yaoxu Song, Shuwen Lu 等ASPLOS 2023 · 被引用 25 次
- SYCL++: A Unified Programming Framework for Heterogeneous Supercomputers at ScaleZitao Shen, Yuyang Jin, Kinman Lei, Zixuan Ma 等HPDC 2026
- UniSparse: An Intermediate Language for General Sparse Format CustomizationJie Liu, Zhongyuan Zhao, Zijian Ding, Benjamin Brock 等OOPSLA 2024 · 被引用 7 次
- NutCracker: A Compilation Framework for Hybrid DPU ArchitecturesYihan Yang, Haifeng Sun, Antoine Kaufmann, Jialin LiEuroSys 2026 · 被引用 2 次
- CINM (Cinnamon): A Compilation Infrastructure for Heterogeneous Compute In-Memory and Compute Near-Memory ParadigmsAsif Ali Khan, Hamid Farzaneh, Karl Friedrich Alexander Friebel, Clément Fournier 等ASPLOS 2024 · 被引用 7 次
