QuCo: Efficient and Flexible Hardware-Driven Automatic Configuration of Tile Transfers in GPUs
Nicolás Meseguer, Daoxuan Xu, Yifan Sun, Michael Pellauer, José L. Abellán, Manuel E. Acacio
Abstract
The growing complexity and parallelism demands of modern GPU workloads have driven architectural innovations toward asynchronous tile transfers (ATTs) to overlap computation and data movement. While ATT units such as the NVIDIA's Tensor Memory Accelerator (TMA) introduce high-throughput memory transfers, programmers must deal with wavefront specialization, select tile sizes, queue slots, and synchronization primitives, all of which are hardware-specific and workloaddependent. Existing GPU libraries fall short-offering limited ATT support and configurability-so developers still resort to manual exploration of this vast parameter space, which is laborious, error-prone, and fundamentally limits performance portability across GPUs. In this work, we present QuCo (Queue Configurator), a single lightweight hardware unit embedded in the GPU that fully automates the ATT configuration process. Inspired by Blackwell GPU design, QuCo includes a compact RISC-V processor, small memory structures for instructions and data, and a GPU Specification Table (GST) storing key architectural parameters. Using the GST and workload characteristics, along with built-in heuristics, QuCo computes optimal queue configurations at kernel launch. This relieves the programmer of the tedious, time-consuming task of tuning and offline profiling, while simultaneously increasing post-compilation performance portability.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 0abe785c-70c6-4d4e-a017-cf127e938f1aRelated papers
- Task-Based Tensor Computations on Modern GPUsRohan Yadav, Michael Garland, Alex Aiken, Michael BauerPLDI 2025 · 5 citations
- Optimal Software Pipelining and Warp Specialization for Tensor Core GPUsRupanshu Soi, Rohan Yadav, Fredrik Kjolstad, Alex Aiken et al.OSDI 2026 · 9 citations
- Cooperative Warp Execution in Tensor Core for RISC-V GPGPUAbubakr Nada, Giuseppe Maria Sarda, Erwan LenormandHPCA 2025 · 3 citations
- WASP: Exploiting GPU Pipeline Parallelism with Hardware-Accelerated Automatic Warp SpecializationNeal Clayton Crago, Sana Damani, Karthikeyan Sankaralingam, Stephen W. KecklerHPCA 2024 · 13 citations
- LIBRA: Memory Bandwidth- and Locality-Aware Parallel Tile RenderingAurora Tomás, Juan L. Aragón, Joan-Manuel Parcerisa, Antonio GonzálezMICRO 2024 · 1 citation
