Dynamic Scheduling for AI Accelerators via TISA
Guanghui Song, Xiaoqiang Dan, Chengke Wang, Fei Liu, Wenyuan Lv, Zhongzhou Jiang, Jianjian Guan, Teng Lu, Lin Tao, Cheng Li, Weixing Pan, Wei Huang
摘要
Modern AI accelerators suffer from low utilization because static compile-time schedules cannot adapt to runtime variability or coordinate heterogeneous units effectively. This paper presents a semantics-aware dynamic tile scheduling framework that restores the missing runtime semantics required for adaptive execution. It co-designs three synergistic components, including a semantics-preserving compiler that maintains operator boundaries and dependency types through lowering, a tile-level instruction set (TISA) that encodes typed dependencies, resource intents, and tile-level memory ranges, and a conflictaware runtime scheduler that uses these semantics to dynamically reorder tiles, resolve contention, and overlap execution across tensor, vector, and DMA units. This design unifies software semantics and hardware scheduling, enabling cross-operator and cross-iteration parallelism beyond static approaches. Across ResNet50, BERT, GPT-J, and LLaMA2, our work achieves 1.52-1.92× speedups over the baseline, delivering 1.14-1.63× additional improvement over strong static tile-level pipeline scheduling; on FlashAttention-3 (head dim 128), it improves hardware utilization by 26.4% versus the state-of-the-art H100 implementation. Ablation studies further show that semantics preservation alone yields gains, confirming the independent value of restoring scheduling information to runtime.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- PipeThreader: Software-Defined Pipelining for Efficient DNN ExecutionYu Cheng, Lei Wang, Yining Shi, Yuqing Xia 等OSDI 2025 · 被引用 9 次
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionJay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar 等NeurIPS 2024 · 被引用 727 次
- Trinity: Three-Dimensional Tensor Program Optimization via Tile-level Equality SaturationJaehyeong Park, Youngchan Kim, Haechan An, Gieun Jeong 等ASPLOS 2026
- TileLang: Bridge Programmability and Performance in Modern Neural KernelsLei Wang, Yu Cheng, Yining Shi, Zhiwen Mo 等ICLR 2026
- HYTE: Flexible Tiling for Sparse Accelerators via Hybrid Static-Dynamic ApproachesXintong Li, Zhiyao Li, Mingyu GaoISCA 2025 · 被引用 2 次
