Lune

ISCA2024Top-tier venue

A New Formulation of Neural Data Prefetching

Quang Duong, Akanksha Jain, Calvin Lin

2024Year
16Citations
4Top-tier citations

Abstract

Temporal data prefetchers have the potential to produce significant performance gains by prefetching irregular data streams. Recent work has introduced a neural model for temporal prefetching that outperforms practical table-based temporal prefetchers, but the large storage and latency costs, along with the inability to generalize to memory addresses outside of the training dataset, prevent such a neural network from seeing any practical use in hardware.

In this paper, we reformulate the temporal prefetching prediction problem so that neural solutions to it are more amenable for hardware deployment. Our key insight is that while temporal prefetchers typically assume that each address can be followed by any possible successor, there are empirically only a few successors for each address. Utilizing this insight, we introduce a new abstraction of memory addresses, and we show how this abstraction enables the design of a much more efficient neural prefetcher.

Our new prefetcher, Twilight, improves upon the previous state-of-the-art neural prefetcher, Voyager, in multiple dimensions: It reduces latency by 988×, shrinks storage by 10.8×, achieves 4% more speedup on a mix of irregular SPEC 2006, SPEC 2017, and GAP benchmarks, and is capable of predicting new temporal correlations not present in the training data. Twilight outperforms idealized versions of the non-neural temporal prefetchers STMS by 12.2% and Domino by 8.5%. While Twilight is still not practical, T-LITE, a slimmed-down version of Twilight that can prefetch across different program runs, further reduces latency and storage (1421× faster and 142× smaller than Voyager), matches Voyager's performance and outperforms the practical non-neural Triage prefetcher by 5.9%.

Data prefetchers are vital mechanisms for hiding the long latencies of memory accesses. While most modern hardware prefetchers identify strides or spatial footprints to target regular or spatial access patterns, this paper focuses on temporal prefetching, a type of irregular data prefetching that identifies pairs of addresses that are temporally correlated. For instance, if address X is often followed by address Y, then X and Y are correlated, and a load of X can serve as a trigger to prefetch Y. Since these correlations can be found between any two addresses X and Y, temporal prefetchers can eliminate cache misses from any arbitrary repeated memory access stream.

Recent work by Shi, et al. [41] shows that ML-based temporal prefetchers like Voyager provide significant headroom over idealized versions of table-based temporal prefetchers, achieving an extra 13% speedup. 1 Unfortunately, several limitations prevent Voyager from being realized in practice.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 66ff6720-401f-4cfe-9a12-aa1b04416165

Cited by top-tier papers4

Ask how each one uses it

Builds on11

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines