TENET-v2: Applying Relation-Centric Notation to Model and Optimize Data Swizzle in the Cache of Modern NPU
Hanyu Zhang, Fangxu Guo, Liqiang Lu, Long Wang, Yunfei Du, Zhe Wang, Jinghan Zhang, Jie Zhang, Chenli Xue, Chengpeng Wu, Ziyi Zhang, Yun Liang
Abstract
Swizzle is a data access pattern optimization technique by reorganizing the execution order of computational tasks to improve the cache locality in modern NPUs. Existing analysis and optimization techniques lack support for swizzleaware modeling on NPUs and fail to effectively capture cache behavior across diverse swizzle configurations. To this end, we propose TENET-v2, a framework for modeling and optimizing swizzle. We introduce a relation-centric notation to characterize different cache access patterns, thus exploring wider swizzle space. Then, we propose a hybrid performance model for cache analysis. The proposed performance model uses an analytical approach to quantify cache miss behavior under unsaturated cache conditions (non-saturated misses), and employs a simulation method combined with an early exiting mechanism to rapidly model cache behavior under saturated cache conditions (saturated misses). Experimental evaluations demonstrate that TENET-v2 achieves an average absolute error of 1.05 % in read hit rate compared to real-world hardware. Evaluation on a variety of DNNs shows that TENET-v2 outperforms existing tensor program optimizers by up toon A100 GPUs. We also demonstrate NPU cache size optimization based on TENET-v2.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- TENET: A Framework for Modeling Tensor Dataflow Based on Relation-centric NotationLiqiang Lu, Naiqing Guan, Yuyue Wang, Liancheng Jia et al.ISCA 2021 · 82 citations
- TD-NUCA: Runtime Driven Management of NUCA Caches in Task Dataflow Programming ModelsPaul Caheny, Lluc Alvarez, Marc Casas, Miquel MoretóSC 2022 · 4 citations
- CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUsTianhao Cai, Liang Wang, Limin Xiao, Meng Han et al.DAC 2025
- NVR: Vector Runahead on NPUs for Sparse Memory AccessHui Wang, Zhengpeng Zhao, Jing Wang, Yushu Du et al.DAC 2025 · 1 citation
- MISAAL: Synthesis-Based Automatic Generation of Efficient and Retargetable Semantics-Driven OptimizationsAbdul Rafae Noor, Dhruv Baronia, Akash Kothari, Muchen Xu et al.PLDI 2025 · 2 citations
