SC2022Top-tier venue
TD-NUCA: Runtime Driven Management of NUCA Caches in Task Dataflow Programming Models
Paul Caheny, Lluc Alvarez, Marc Casas, Miquel Moretó
Abstract
In high performance processors, the design of on-chip memory hierarchies is crucial for performance and energy efficiency. Current processors rely on large shared Non-Uniform Cache Architectures (NUCA) to improve performance and reduce data movement. Multiple solutions exploit information available at the microarchitecture level or in the operating system to optimize NUCA performance. However, existing methods have not taken advantage of the information captured by task dataflow programming models to guide the management of NUCA caches. In this paper we propose TD-NUCA, a hardware/software co-designed approach that leverages information present in the run-time system of task dataflow programming models to efficiently manage NUCA caches. TD-NUCA identifies the data access and reuse patterns of parallel applications in the runtime system and guides the operation of the NUCA caches in the hardware. As a result, TD-NUCA achieves a 1.18x average speedup over the baseline S-NUCA while requiring only 0.62x the data movement.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get bfc7d7c8-e0f7-4826-bd20-836b6ddb6a82Cited by top-tier papers2
- Stream-Based Data Placement for Near-Data Processing with Extended MemoryYiwei Li, Boyu Tian, Yi Ren, Mingyu GaoMICRO 2024 · 5 citations
- CPElide: Efficient Multi-Chiplet GPU Implicit SynchronizationPreyesh Dalmia, Rajesh Shashi Kumar, Matthew D. SinclairMICRO 2024 · 3 citations
Related papers
- NUPEA: Optimizing Critical Loads on Spatial Dataflow Architectures via Non-Uniform Processing-Element AccessSouradip Ghosh, Graham Gobieski, Keyi Zhang, Brandon Lucia et al.ISCA 2025 · 1 citation
- TaskStream: accelerating task-parallel workloads by recovering program structureVidushi Dadu, Tony NowatzkiASPLOS 2022 · 23 citations
- TENET-v2: Applying Relation-Centric Notation to Model and Optimize Data Swizzle in the Cache of Modern NPUHanyu Zhang, Fangxu Guo, Liqiang Lu, Long Wang et al.HPCA 2026
- Streaming Task Graph Scheduling for Dataflow ArchitecturesTiziano De Matteis, Lukas Gianinazzi, Johannes de Fine Licht, Torsten HoeflerHPDC 2023 · 3 citations
- Sigma: Compiling Einstein Summations to Locality-Aware DataflowTian Zhao, Alexander Rucker, Kunle OlukotunASPLOS 2023 · 3 citations
