T4: Compiling Sequential Code for Effective Speculative Parallelization in Hardware
Victor A. Ying, Mark C. Jeffrey, Daniel Sánchez
Abstract
Multicores are now ubiquitous, but programmers still write sequential code. Speculative parallelization is an enticing approach to parallelize code while retaining the ease of sequential programming, making parallelism pervasive. However, prior speculative parallelizing compilers and architectures achieved limited speedups due to high costs of recovering from misspeculation and hardware scalability bottlenecks.
We present T4, a parallelizing compiler that successfully leverages recent hardware features for speculative execution, which present new opportunities and challenges for automatic parallelization. T4 transforms sequential programs into trees of tiny timestamped tasks. T4 introduces novel compiler techniques to expose parallelism aggressively across the entire program, breaking applications into tiny tasks of tens of instructions each. Task trees unfold their branches in parallel to enable high taskspawn throughput while exploiting selective aborts to recover from misspeculation cheaply. T4 exploits parallelism across function calls, loops, and loop nests; performs new transformations to reduce task spawn costs and avoid false sharing; and exploits data locality among fine-grain tasks. As a result, T4 scales several hard-to-parallelize SPEC CPU2006 benchmarks to tens of cores, on which prior work attained little or no speedup.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d371faa3-ee95-4111-81f5-52203d36f81aCited by top-tier papers3
- TaskStream: accelerating task-parallel workloads by recovering program structureVidushi Dadu, Tony NowatzkiASPLOS 2022 · 23 citations
- Taming the Zoo: The Unified GraphIt Compiler Framework for Novel ArchitecturesAjay Brahmakshatriya, Emily Furst, Victor A. Ying, Claire Hsu et al.ISCA 2021 · 12 citations
- A scalable architecture for reprioritizing ordered parallelismGilead Posluns, Yan Zhu, Guowei Zhang, Mark C. JeffreyISCA 2022 · 8 citations
Builds on2
Related papers
- LoopFrog: In-Core Hint-Based Loop ParallelizationMárton Erdos, Utpal Bora, Akshay Bhosale, Bob Lytton et al.MICRO 2025 · 1 citation
- Symbiotic Task Scheduling and Data PrefetchingGilead Posluns, Mark C. JeffreyMICRO 2025 · 1 citation
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Phloem: Automatic Acceleration of Irregular Applications with Fine-Grain Pipeline ParallelismQuan M. Nguyen, Daniel SánchezHPCA 2023 · 7 citations
- Attention-Level SpeculationJack Cai, Ammar Vora, Randolph Zhang, Mark O'Connor et al.ICML 2025
