Phloem: Automatic Acceleration of Irregular Applications with Fine-Grain Pipeline Parallelism
Quan M. Nguyen, Daniel Sánchez
Abstract
Irregular applications are increasingly common in diverse domains, like graph analytics and sparse linear algebra. Accelerating these applications is challenging because of their unpredictable data reuse and control flow. Recent work has proposed hardware support for fine-grain pipeline parallelism, hiding long latencies by decoupling irregular applications into pipeline stages. However, this prior work requires programmers to manually decouple applications. This tedious and error-prone process limits the usefulness of such architectural support.
We address this problem with Phloem, a compiler that automatically discovers and exploits pipeline parallelism in irregular applications. Prior compilers for pipeline parallelism target regular applications, which contain simple pipeline stages with known latencies and fixed buffering needs. Designing Phloem to target irregular applications, where these properties do not hold, requires treating their unique challenges as first-class considerations throughout its design. Phloem breaks down this complex transformation into a series of simple passes that together encode the insights that have been previously applied by hand, producing code that targets architectures with support for queue-based communication.
We evaluate Phloem by generating efficient pipelines on a variety of irregular applications. Phloem's contributions improve performance by 1.7× on average, approaching (and sometimes exceeding) the performance of manually optimized pipelineparallel code. These results show that, for the first time, automatic parallelization for irregular applications is not only feasible, but also profitable.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef4cbba6-5f4e-47da-93f4-e3fd3f0bd775Cited by top-tier papers4
- ASPEN: Breaking Operator Barriers for Efficient Parallelization of Deep Neural NetworksJongseok Park, Kyungmin Bin, Gibum Park, Sangtae Ha et al.NeurIPS 2023 · 7 citations
- Ripple: Asynchronous Programming for Spatial Dataflow ArchitecturesSouradip Ghosh, Yufei Shi, Brandon Lucia, Nathan BeckmannPLDI 2025 · 4 citations
- NUPEA: Optimizing Critical Loads on Spatial Dataflow Architectures via Non-Uniform Processing-Element AccessSouradip Ghosh, Graham Gobieski, Keyi Zhang, Brandon Lucia et al.ISCA 2025 · 1 citation
- Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML FusionArash Nasr-Esfahany, Mohammad Alizadeh, Victor Lee, Hanna Alam et al.ISCA 2025 · 1 citation
Builds on8
- Prodigy: Improving the Memory Latency of Data-Indirect Irregular Workloads Using Hardware-Software Co-DesignNishil Talati, Kyle May, Armand Behroozi, Yichen Yang et al.HPCA 2021 · 62 citations
- Fifer: Practical Acceleration of Irregular Applications on Reconfigurable ArchitecturesQuan M. Nguyen, Daniel SánchezMICRO 2021 · 60 citations
- SARA: Scaling a Reconfigurable Dataflow AcceleratorYaqi Zhang, Nathan Zhang, Tian Zhao, Matt Vilim et al.ISCA 2021 · 58 citations
- A compiler infrastructure for accelerator generatorsRachit Nigam, Samuel Thomas, Zhijing Li, Adrian SampsonASPLOS 2021 · 54 citations
- Type-directed scheduling of streaming acceleratorsDavid Durst, Matthew Feldman, Dillon Huff, David Akeley et al.PLDI 2020 · 49 citations
Related papers
- Pipette: Improving Core Utilization on Irregular Applications through Intra-Core Pipeline ParallelismQuan M. Nguyen, Daniel SánchezMICRO 2020 · 28 citations
- SpZip: Architectural Support for Effective Data Compression In Irregular ApplicationsYifan Yang, Joel S. Emer, Daniel SánchezISCA 2021 · 31 citations
- WASP: Exploiting GPU Pipeline Parallelism with Hardware-Accelerated Automatic Warp SpecializationNeal Clayton Crago, Sana Damani, Karthikeyan Sankaralingam, Stephen W. KecklerHPCA 2024 · 13 citations
- Sigma: Compiling Einstein Summations to Locality-Aware DataflowTian Zhao, Alexander Rucker, Kunle OlukotunASPLOS 2023 · 3 citations
- PANORAMA: divide-and-conquer approach for mapping complex loop kernels on CGRADhananjaya Wijerathne, Zhaoying Li, Thilini Kaushalya Bandara, Tulika MitraDAC 2022 · 18 citations
