Ghost Threading: Helper-Thread Prefetching for Real Systems
Yuxin Guo, Akshay Bhosale, Utpal Bora, Alexandra W. Chadwick, Márton Erdos, Giacomo Gabrielli, Timothy M. Jones
Abstract
Memory latency is the bottleneck for many modern workloads.One popular solution from literature to handle this is helper threading, a technique that issues light-weight prefetching helper thread(s) extracted from the original application to bring data into the cache before the main thread uses it, hiding the long memory latency.Although prior work has reported promising results, many schemes are not available on real systems as they require hardware support for satisfying performance improvements.To address this, we present Ghost Threading, a software-only helper-thread prefetching solution, which issues helper threads on idle Simultaneous Multithreading (SMT) contexts.The key challenge of prefetching is timeliness: data should not arrive too early or too late.Unlike prior work relying on proposed extra hardware synchronization or expensive OS synchronization, we develop a novel inter-thread synchronization approach based on an instruction supported by commercial processors, which enables cheap throttling of the helper thread.This ensures that it gets far enough ahead of the main thread, but not too far, for it to perform timely prefetching of data.We evaluate Ghost Threading against state-of-the-art techniques on a modern Intel processor with memory-intensive workloads selected from graph analysis, database, and HPC domains.On an idle server, Ghost Threading achieves 1.33× geometric mean speedup over the baseline, 1.25× and 1.11× over state-of-the-art software prefetching and parallelization techniques, respectively.We also show that Ghost Threading maintains these benefits on a busy server where the memory bandwidth pressure is higher.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Reducing Load Latency with Cache Level PredictionMajid Jalili, Mattan ErezHPCA 2022 · 17 citations
- GhOST: a GPU Out-of-Order Scheduling Technique for Stall ReductionIshita Chaturvedi, Bhargav Reddy Godala, Yucan Wu, Ziyang Xu et al.ISCA 2024 · 9 citations
- Harvesting Memory-bound CPU Stall Cycles in Software with MSHZhihong Luo, Sam Son, Sylvia Ratnasamy, Scott ShenkerOSDI 2024 · 5 citations
- SMTcheck: Accurate SMT Interference Prediction to Improve Scheduling Efficiency in DatacentersSanghyun Kim, Jinhyeok Oh, Taehun Kim, Gyutae Kim et al.HPCA 2026
- Decoupled Vector RunaheadAjeya Naithani, Jaime Roelandts, Sam Ainsworth, Timothy M. Jones et al.MICRO 2023 · 15 citations
