TD-NUCA: Runtime Driven Management of NUCA Caches in Task Dataflow Programming Models
Paul Caheny, Lluc Alvarez, Marc Casas, Miquel Moretó
摘要
In high performance processors, the design of on-chip memory hierarchies is crucial for performance and energy efficiency. Current processors rely on large shared Non-Uniform Cache Architectures (NUCA) to improve performance and reduce data movement. Multiple solutions exploit information available at the microarchitecture level or in the operating system to optimize NUCA performance. However, existing methods have not taken advantage of the information captured by task dataflow programming models to guide the management of NUCA caches. In this paper we propose TD-NUCA, a hardware/software co-designed approach that leverages information present in the run-time system of task dataflow programming models to efficiently manage NUCA caches. TD-NUCA identifies the data access and reuse patterns of parallel applications in the runtime system and guides the operation of the NUCA caches in the hardware. As a result, TD-NUCA achieves a 1.18x average speedup over the baseline S-NUCA while requiring only 0.62x the data movement.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Stream-Based Data Placement for Near-Data Processing with Extended MemoryYiwei Li, Boyu Tian, Yi Ren, Mingyu GaoMICRO 2024 · 被引用 5 次
- CPElide: Efficient Multi-Chiplet GPU Implicit SynchronizationPreyesh Dalmia, Rajesh Shashi Kumar, Matthew D. SinclairMICRO 2024 · 被引用 3 次
相关 Paper
- NUPEA: Optimizing Critical Loads on Spatial Dataflow Architectures via Non-Uniform Processing-Element AccessSouradip Ghosh, Graham Gobieski, Keyi Zhang, Brandon Lucia 等ISCA 2025 · 被引用 1 次
- TaskStream: accelerating task-parallel workloads by recovering program structureVidushi Dadu, Tony NowatzkiASPLOS 2022 · 被引用 23 次
- TENET-v2: Applying Relation-Centric Notation to Model and Optimize Data Swizzle in the Cache of Modern NPUHanyu Zhang, Fangxu Guo, Liqiang Lu, Long Wang 等HPCA 2026
- Streaming Task Graph Scheduling for Dataflow ArchitecturesTiziano De Matteis, Lukas Gianinazzi, Johannes de Fine Licht, Torsten HoeflerHPDC 2023 · 被引用 3 次
- Sigma: Compiling Einstein Summations to Locality-Aware DataflowTian Zhao, Alexander Rucker, Kunle OlukotunASPLOS 2023 · 被引用 3 次
