ThreadFuser: A SIMT Analysis Framework for MIMD Programs
Ahmad Alawneh, Ni Kang, Mahmoud Khairy, Timothy G. Rogers
Abstract
The broad usage of accelerators, such as GPUs, faces two important challenges. Developing code for a new accelerator is expensive and unpredictable. Porting large parallel programs from Multiple Instruction Multiple Data (MIMD) CPUs to Single Instruction Multiple Thread (SIMT) GPUs involves significant effort that may or may not result in improved performance versus the CPU. This high activation energy to create new workloads introduces the second challenge: architects and systems researchers lack a diverse SIMT codebase to study new designs.
To tackle these challenges, we introduce ThreadFuser, an analysis framework that efficiently and accurately predicts the performance of any pre-written MIMD program on SIMT hardware. ThreadFuser conducts thorough control and data flow analysis on dynamic CPU program traces, determining the impact of lock-step execution on CPU binaries. Thread-Fuser efficiently delivers accurate reports on a MIMD program's divergence and synchronization characteristics. Moreover, ThreadFuser seamlessly integrates with state-of-the-art GPU simulators to conduct detailed analyses and produce fine-grained performance measurements.
We evaluate ThreadFuser on a diverse set of 36 CPU workloads, demonstrating the potential and challenges of executing MIMD code on a SIMT machine. We demonstrate ThreadFuser's potential to inform software development decisions and open new areas to explore in data-parallel hardware design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on3
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 366 citations
- ndzip-gpu: efficient lossless compression of scientific floating-point data on GPUsFabian Knorr, Peter Thoman, Thomas FahringerSC 2021 · 29 citations
- SIMR: Single Instruction Multiple Request Processing for Energy-Efficient Data Center MicroservicesMahmoud Khairy, Ahmad Alawneh, Aaron Barnes, Timothy G. RogersMICRO 2022 · 8 citations
Related papers
- FlipFlop: A Static Analysis-based Energy Optimization Framework for GPU KernelsSaurabhsingh Rajput, Alexander Brandt, Vadim Elisseev, Tushar SharmaICSE 2026
- Principal Kernel Analysis: A Tractable Methodology to Simulate Scaled GPU WorkloadsCesar Avalos Baddouh, Mahmoud Khairy, Roland N. Green, Mathias Payer et al.MICRO 2021 · 26 citations
- A Modular Static Cost Analysis for GPU Warp-Level ParallelismGregory Blike, Hannah Zicarelli, Udaya Sathiyamoorthy, Julien Lange et al.POPL 2026 · 1 citation
- MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive OperatorsZheng Zhang, Donglin Yang, Xiaobo Zhou, Dazhao ChengSC 2024 · 9 citations
- Fuzzing Open-Source GPU Hardware with SIMT Program GenerationZibo Gao, Jie Wang, Qihang Zhou, Lixiao Shan et al.USENIX Security 2026
