Dynamic Scheduling for AI Accelerators via TISA
Guanghui Song, Xiaoqiang Dan, Chengke Wang, Fei Liu, Wenyuan Lv, Zhongzhou Jiang, Jianjian Guan, Teng Lu, Lin Tao, Cheng Li, Weixing Pan, Wei Huang
Abstract
Modern AI accelerators suffer from low utilization because static compile-time schedules cannot adapt to runtime variability or coordinate heterogeneous units effectively. This paper presents a semantics-aware dynamic tile scheduling framework that restores the missing runtime semantics required for adaptive execution. It co-designs three synergistic components, including a semantics-preserving compiler that maintains operator boundaries and dependency types through lowering, a tile-level instruction set (TISA) that encodes typed dependencies, resource intents, and tile-level memory ranges, and a conflictaware runtime scheduler that uses these semantics to dynamically reorder tiles, resolve contention, and overlap execution across tensor, vector, and DMA units. This design unifies software semantics and hardware scheduling, enabling cross-operator and cross-iteration parallelism beyond static approaches. Across ResNet50, BERT, GPT-J, and LLaMA2, our work achieves 1.52-1.92× speedups over the baseline, delivering 1.14-1.63× additional improvement over strong static tile-level pipeline scheduling; on FlashAttention-3 (head dim 128), it improves hardware utilization by 26.4% versus the state-of-the-art H100 implementation. Ablation studies further show that semantics preservation alone yields gains, confirming the independent value of restoring scheduling information to runtime.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- PipeThreader: Software-Defined Pipelining for Efficient DNN ExecutionYu Cheng, Lei Wang, Yining Shi, Yuqing Xia et al.OSDI 2025 · 9 citations
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionJay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar et al.NeurIPS 2024 · 727 citations
- Trinity: Three-Dimensional Tensor Program Optimization via Tile-level Equality SaturationJaehyeong Park, Youngchan Kim, Haechan An, Gieun Jeong et al.ASPLOS 2026
- TileLang: Bridge Programmability and Performance in Modern Neural KernelsLei Wang, Yu Cheng, Yining Shi, Zhiwen Mo et al.ICLR 2026
- HYTE: Flexible Tiling for Sparse Accelerators via Hybrid Static-Dynamic ApproachesXintong Li, Zhiyao Li, Mingyu GaoISCA 2025 · 2 citations
