Uncovering Nested Data Parallelism and Data Reuse in DNN Computation with FractalTensor
Siran Liu, Chengxiang Qi, Ying Cao, Chao Yang, Weifang Hu, Xuanhua Shi, Fan Yang, Mao Yang
Abstract
To speed up computation, deep neural networks (DNNs) usually rely on highly optimized tensor operators. Despite the effectiveness, tensor operators are often defined empirically with ad hoc semantics. This hinders the analysis and optimization across operator boundaries. FractalTensor is a programming framework that addresses this challenge. At the core, FractalTensor is a nested list-based abstract data type (ADT), where each element is a tensor with static shape or another FractalTensor (i.e., nested). DNNs are then de-fined by high-order array compute operators like map/reduce/scan and array access operators like window/stride on FractalTensor. This new way of DNN definition explicitly exposes nested data parallelism and fine-grained data access patterns, opening new opportunities for whole program analysis and optimization. To exploit these opportunities, from the FractalTensor-based code the compiler extracts a nested multi-dimensional dataflow graph called Extended Task Dependence Graph (ETDG), which provides a holistic view of data dependency across different granularity. The ETDG is then transformed into an efficient implementation through graph coarsening, data reordering, and access materialization. Evaluation on six representative DNNs like RNN and FlashAttention on NVIDIA A100 shows that Fractal-Tensor achieves speedup by up to 5.45x and 2.14x on average through a unified solution for diverse optimizations.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6548b844-8dec-40b5-af0e-28cb1babfc12Cited by top-tier papers4
- PipeThreader: Software-Defined Pipelining for Efficient DNN ExecutionYu Cheng, Lei Wang, Yining Shi, Yuqing Xia et al.OSDI 2025 · 9 citations
- PluS: Highly Efficient and Expandable ML Compiler with Pluggable Graph SchedulesRuofan Wu, Zhen Zheng, Feng Zhang, Chuanjie Liu et al.USENIX ATC 2025 · 5 citations
- TrainVerify: Equivalence-Based Verification for Distributed LLM TrainingYunchi Lu, Youshan Miao, Cheng Tan, Peng Huang et al.SOSP 2025 · 1 citation
- Tempo: Compiled Dynamic Deep Learning with Symbolic Dependence GraphsPedro F. Silvestre, Peter R. PietzuchSOSP 2025
Related papers
- FreeTensor: a free-form DSL with holistic optimizations for irregular tensor programsShizhi Tang, Jidong Zhai, Haojie Wang, Lin Jiang et al.PLDI 2022 · 16 citations
- Neptune: Advanced ML Operator Fusion for Locality and Parallelism on GPUsYifan Zhao, Egan Johnson, Prasanth Chatarasi, Vikram S. Adve et al.PLDI 2026 · 1 citation
- Optimizing Deep Learning Inference via Global Analysis and Tensor ExpressionsChunwei Xia, Jiacheng Zhao, Qianqi Sun, Zheng Wang et al.ASPLOS 2024 · 14 citations
- Fractal: Joint Multi-Level Sparse Pattern Tuning of Accuracy and Performance for DNN PruningYue Guan, Changming Yu, Yangjie Zhou, Jingwen Leng et al.ASPLOS 2024 · 6 citations
- Optimizing Deep Learning Inference Efficiency through Block Dependency AnalysisZhanyuan Di, Leping Wang, En Shao, Zhaojia Ma et al.ASPLOS 2025 · 2 citations
