ASPEN: Breaking Operator Barriers for Efficient Parallelization of Deep Neural Networks
Jongseok Park, Kyungmin Bin, Gibum Park, Sangtae Ha, Kyunghan Lee
摘要
Modern Deep Neural Network (DNN) frameworks use tensor operators as the main building blocks of DNNs. However, we observe that operator-based construction of DNNs incurs significant drawbacks in parallelism in the form of synchronization barriers. Synchronization barriers of operators confine the scope of parallel computation to each operator and obscure the rich parallel computation opportunities that exist across operators. To this end, we present ASPEN, a novel parallel computation solution for DNNs that allows fine-grained dynamic execution of DNNs, which (1) removes the operator barriers and expresses DNNs in dataflow graphs of fine-grained tiles to expose the parallel computation opportunities across operators, and (2) exploits these opportunities by dynamically locating and scheduling them in runtime. This novel approach of ASPEN enables opportunistic parallelism, a new class of parallelism for DNNs that is unavailable in the existing operator-based approaches. ASPEN also achieves high resource utilization and memory reuse by letting each resource asynchronously traverse depthwise in the DNN graph to its full computing potential. We provide challenges and solutions to our approach and show that our proof-of-concept implementation of ASPEN on CPU shows exceptional performance, outperforming state-of-the-art inference systems of TorchScript and TVM by up to 3.2× and 4.3×, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- NanoFlow: Towards Optimal Large Language Model Serving ThroughputKan Zhu, Yufei Gao, Yilong Zhao, Liangyu Zhao 等OSDI 2025 · 被引用 92 次
- Mirage: A Multi-Level Superoptimizer for Tensor ProgramsMengdi Wu, Xinhao Cheng, Shengyu Liu, Chunan Shi 等OSDI 2025 · 被引用 49 次
- LoRAFusion: Efficient LoRA Fine-Tuning for LLMsZhanda Zhu, Qidong Su, Yaoyao Ding, Kevin Song 等EuroSys 2026 · 被引用 2 次
- Trinity: Three-Dimensional Tensor Program Optimization via Tile-level Equality SaturationJaehyeong Park, Youngchan Kim, Haechan An, Gieun Jeong 等ASPLOS 2026
它引用的顶会 Paper12
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep LearningAurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger 等OSDI 2021 · 被引用 258 次
- DAPPLE: a pipelined data parallel approach for training large modelsShiqing Fan, Yi Rong, Chen Meng, Zongyan Cao 等PPoPP 2021 · 被引用 224 次
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue 等OSDI 2020 · 被引用 192 次
- Unity: Accelerating DNN Training Through Joint Optimization of Algebraic Transformations and ParallelizationColin Unger, Zhihao Jia, Wei Wu, Sina Lin 等OSDI 2022 · 被引用 105 次
- Nimble: Lightweight and Parallel GPU Task Scheduling for Deep LearningWoosuk Kwon, Gyeong-In Yu, Eunji Jeong, Byung-Gon ChunNeurIPS 2020 · 被引用 102 次
相关 Paper
- Optimizing Deep Learning Inference via Global Analysis and Tensor ExpressionsChunwei Xia, Jiacheng Zhao, Qianqi Sun, Zheng Wang 等ASPLOS 2024 · 被引用 14 次
- Uncovering Nested Data Parallelism and Data Reuse in DNN Computation with FractalTensorSiran Liu, Chengxiang Qi, Ying Cao, Chao Yang 等SOSP 2024 · 被引用 1 次
- GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline ParallelismByungsoo Jeon, Mengdi Wu, Shiyi Cao, Sunghyun Kim 等ASPLOS 2025 · 被引用 10 次
- FreeTensor: a free-form DSL with holistic optimizations for irregular tensor programsShizhi Tang, Jidong Zhai, Haojie Wang, Lin Jiang 等PLDI 2022 · 被引用 16 次
- SoD2: Statically Optimizing Dynamic Deep Neural Network ExecutionWei Niu, Gagan Agrawal, Bin RenASPLOS 2024 · 被引用 6 次
