T4: Compiling Sequential Code for Effective Speculative Parallelization in Hardware
Victor A. Ying, Mark C. Jeffrey, Daniel Sánchez
摘要
Multicores are now ubiquitous, but programmers still write sequential code. Speculative parallelization is an enticing approach to parallelize code while retaining the ease of sequential programming, making parallelism pervasive. However, prior speculative parallelizing compilers and architectures achieved limited speedups due to high costs of recovering from misspeculation and hardware scalability bottlenecks.
We present T4, a parallelizing compiler that successfully leverages recent hardware features for speculative execution, which present new opportunities and challenges for automatic parallelization. T4 transforms sequential programs into trees of tiny timestamped tasks. T4 introduces novel compiler techniques to expose parallelism aggressively across the entire program, breaking applications into tiny tasks of tens of instructions each. Task trees unfold their branches in parallel to enable high taskspawn throughput while exploiting selective aborts to recover from misspeculation cheaply. T4 exploits parallelism across function calls, loops, and loop nests; performs new transformations to reduce task spawn costs and avoid false sharing; and exploits data locality among fine-grain tasks. As a result, T4 scales several hard-to-parallelize SPEC CPU2006 benchmarks to tens of cores, on which prior work attained little or no speedup.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- TaskStream: accelerating task-parallel workloads by recovering program structureVidushi Dadu, Tony NowatzkiASPLOS 2022 · 被引用 23 次
- Taming the Zoo: The Unified GraphIt Compiler Framework for Novel ArchitecturesAjay Brahmakshatriya, Emily Furst, Victor A. Ying, Claire Hsu 等ISCA 2021 · 被引用 12 次
- A scalable architecture for reprioritizing ordered parallelismGilead Posluns, Yan Zhu, Guowei Zhang, Mark C. JeffreyISCA 2022 · 被引用 8 次
它引用的顶会 Paper2
相关 Paper
- LoopFrog: In-Core Hint-Based Loop ParallelizationMárton Erdos, Utpal Bora, Akshay Bhosale, Bob Lytton 等MICRO 2025 · 被引用 1 次
- Symbiotic Task Scheduling and Data PrefetchingGilead Posluns, Mark C. JeffreyMICRO 2025 · 被引用 1 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Phloem: Automatic Acceleration of Irregular Applications with Fine-Grain Pipeline ParallelismQuan M. Nguyen, Daniel SánchezHPCA 2023 · 被引用 7 次
- Attention-Level SpeculationJack Cai, Ammar Vora, Randolph Zhang, Mark O'Connor 等ICML 2025
