TENET-v2: Applying Relation-Centric Notation to Model and Optimize Data Swizzle in the Cache of Modern NPU
Hanyu Zhang, Fangxu Guo, Liqiang Lu, Long Wang, Yunfei Du, Zhe Wang, Jinghan Zhang, Jie Zhang, Chenli Xue, Chengpeng Wu, Ziyi Zhang, Yun Liang
摘要
Swizzle is a data access pattern optimization technique by reorganizing the execution order of computational tasks to improve the cache locality in modern NPUs. Existing analysis and optimization techniques lack support for swizzleaware modeling on NPUs and fail to effectively capture cache behavior across diverse swizzle configurations. To this end, we propose TENET-v2, a framework for modeling and optimizing swizzle. We introduce a relation-centric notation to characterize different cache access patterns, thus exploring wider swizzle space. Then, we propose a hybrid performance model for cache analysis. The proposed performance model uses an analytical approach to quantify cache miss behavior under unsaturated cache conditions (non-saturated misses), and employs a simulation method combined with an early exiting mechanism to rapidly model cache behavior under saturated cache conditions (saturated misses). Experimental evaluations demonstrate that TENET-v2 achieves an average absolute error of 1.05 % in read hit rate compared to real-world hardware. Evaluation on a variety of DNNs shows that TENET-v2 outperforms existing tensor program optimizers by up toon A100 GPUs. We also demonstrate NPU cache size optimization based on TENET-v2.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- TENET: A Framework for Modeling Tensor Dataflow Based on Relation-centric NotationLiqiang Lu, Naiqing Guan, Yuyue Wang, Liancheng Jia 等ISCA 2021 · 被引用 82 次
- TD-NUCA: Runtime Driven Management of NUCA Caches in Task Dataflow Programming ModelsPaul Caheny, Lluc Alvarez, Marc Casas, Miquel MoretóSC 2022 · 被引用 4 次
- CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUsTianhao Cai, Liang Wang, Limin Xiao, Meng Han 等DAC 2025
- NVR: Vector Runahead on NPUs for Sparse Memory AccessHui Wang, Zhengpeng Zhao, Jing Wang, Yushu Du 等DAC 2025 · 被引用 1 次
- MISAAL: Synthesis-Based Automatic Generation of Efficient and Retargetable Semantics-Driven OptimizationsAbdul Rafae Noor, Dhruv Baronia, Akash Kothari, Muchen Xu 等PLDI 2025 · 被引用 2 次
