Uncovering Nested Data Parallelism and Data Reuse in DNN Computation with FractalTensor
Siran Liu, Chengxiang Qi, Ying Cao, Chao Yang, Weifang Hu, Xuanhua Shi, Fan Yang, Mao Yang
摘要
To speed up computation, deep neural networks (DNNs) usually rely on highly optimized tensor operators. Despite the effectiveness, tensor operators are often defined empirically with ad hoc semantics. This hinders the analysis and optimization across operator boundaries. FractalTensor is a programming framework that addresses this challenge. At the core, FractalTensor is a nested list-based abstract data type (ADT), where each element is a tensor with static shape or another FractalTensor (i.e., nested). DNNs are then de-fined by high-order array compute operators like map/reduce/scan and array access operators like window/stride on FractalTensor. This new way of DNN definition explicitly exposes nested data parallelism and fine-grained data access patterns, opening new opportunities for whole program analysis and optimization. To exploit these opportunities, from the FractalTensor-based code the compiler extracts a nested multi-dimensional dataflow graph called Extended Task Dependence Graph (ETDG), which provides a holistic view of data dependency across different granularity. The ETDG is then transformed into an efficient implementation through graph coarsening, data reordering, and access materialization. Evaluation on six representative DNNs like RNN and FlashAttention on NVIDIA A100 shows that Fractal-Tensor achieves speedup by up to 5.45x and 2.14x on average through a unified solution for diverse optimizations.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- PipeThreader: Software-Defined Pipelining for Efficient DNN ExecutionYu Cheng, Lei Wang, Yining Shi, Yuqing Xia 等OSDI 2025 · 被引用 9 次
- PluS: Highly Efficient and Expandable ML Compiler with Pluggable Graph SchedulesRuofan Wu, Zhen Zheng, Feng Zhang, Chuanjie Liu 等USENIX ATC 2025 · 被引用 5 次
- TrainVerify: Equivalence-Based Verification for Distributed LLM TrainingYunchi Lu, Youshan Miao, Cheng Tan, Peng Huang 等SOSP 2025 · 被引用 1 次
- Tempo: Compiled Dynamic Deep Learning with Symbolic Dependence GraphsPedro F. Silvestre, Peter R. PietzuchSOSP 2025
相关 Paper
- FreeTensor: a free-form DSL with holistic optimizations for irregular tensor programsShizhi Tang, Jidong Zhai, Haojie Wang, Lin Jiang 等PLDI 2022 · 被引用 16 次
- Neptune: Advanced ML Operator Fusion for Locality and Parallelism on GPUsYifan Zhao, Egan Johnson, Prasanth Chatarasi, Vikram S. Adve 等PLDI 2026 · 被引用 1 次
- Optimizing Deep Learning Inference via Global Analysis and Tensor ExpressionsChunwei Xia, Jiacheng Zhao, Qianqi Sun, Zheng Wang 等ASPLOS 2024 · 被引用 14 次
- Fractal: Joint Multi-Level Sparse Pattern Tuning of Accuracy and Performance for DNN PruningYue Guan, Changming Yu, Yangjie Zhou, Jingwen Leng 等ASPLOS 2024 · 被引用 6 次
- Optimizing Deep Learning Inference Efficiency through Block Dependency AnalysisZhanyuan Di, Leping Wang, En Shao, Zhaojia Ma 等ASPLOS 2025 · 被引用 2 次
