Divide and Conquer Frontend Bottleneck
Ali Ansari, Pejman Lotfi-Kamran, Hamid Sarbazi-Azad
Abstract
The frontend stalls caused by instruction and BTB misses are a significant source of performance degradation in server processors. Prefetchers are commonly employed to mitigate frontend bottleneck. However, next-line prefetchers, which are available in server processors, are incapable of eliminating a considerable number of L1 instruction misses. Temporal instruction prefetchers, on the other hand, effectively remove most of the instruction and BTB misses but impose significant area overhead.Recently, an old idea of using BTB-directed instruction prefetching is revived to address the limitations of temporal instruction prefetchers. While this approach leads to prefetchers with low area overhead, it requires significant changes to the frontend of a processor. Moreover, as this approach relies on the BTB content for prefetching, BTB misses stall the prefetcher, and likely lead to costly instruction misses. Especially as instruction misses are usually more expensive than BTB misses, the dependence of instruction prefetching to the BTB content is harmful to workloads with very large instruction footprints. Moreover, BTB-directed instruction prefetchers, as proposed in prior work, cannot be applied to variable-length ISAs.In this work, we showcase the harmful effects of making instruction prefetchers depend on the BTB content. Moreover, we divide the frontend bottleneck into three categories and use a divide-and-conquer approach to propose simple and effective solutions for each one. Sequential misses can be covered by an accurate and timely sequential prefetcher named SN4L, a lightweight discontinuity prefetcher named Dis eliminates discontinuity misses, and the BTB misses are reduced by pre-decoding the prefetched blocks. We also discuss how our proposal can be used for variable-length ISAs with low storage overhead. Our proposal, SN4L+ Dis+BTB, imposes the same area overhead as the state-of-the-art BTB-directed prefetcher, and at the same time, outperforms it by 5% on average and up to 16%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd98464e-27f7-4a42-93fa-64f434817a6fCited by top-tier papers9
- Lukewarm serverless functions: characterization and optimizationDavid Schall, Artemiy Margaritov, Dmitrii Ustiugov, Andreas Sandberg et al.ISCA 2022 · 36 citations
- Twig: Profile-Guided BTB Prefetching for Data Center ApplicationsTanvir Ahmed Khan, Nathan Brown, Akshitha Sriraman, Niranjan K. Soundararajan et al.MICRO 2021 · 33 citations
- Cerebros: Evading the RPC Tax in DatacentersArash Pourhabibi Zarandi, Mark Sutherland, Alexandros Daglis, Babak FalsafiMICRO 2021 · 24 citations
- Thermometer: profile-guided btb replacement for data center applicationsShixin Song, Tanvir Ahmed Khan, Sara Mahdizadeh-Shahri, Akshitha Sriraman et al.ISCA 2022 · 23 citations
- PDede: Partitioned, Deduplicated, Delta Branch Target BufferNiranjan K. Soundararajan, Peter Braun, Tanvir Ahmed Khan, Baris Kasikci et al.MICRO 2021 · 22 citations
Related papers
- Morrigan: A Composite Instruction TLB PrefetcherGeorgios Vavouliotis, Lluc Alvarez, Boris Grot, Daniel A. Jiménez et al.MICRO 2021 · 19 citations
- Skia: Exposing Shadow BranchesChrysanthos Pepi, Bhargav Reddy Godala, Krishnam Tibrewala, Gino A. Chacon et al.ASPLOS 2025 · 2 citations
- PDIP: Priority Directed Instruction PrefetchingBhargav Reddy Godala, Sankara Prasad Ramesh, Gilles A. Pokam, Jared Stark et al.ASPLOS 2024 · 17 citations
- Branch Target Buffer OrganizationsArthur Perais, Rami SheikhMICRO 2023 · 8 citations
- Enhancing Instruction Prefetching via Cache and TLB ManagementAlexandre Valentin Jamet, Georgios Vavouliotis, Martí Torrents, Dimitrios Chasapis et al.ISCA 2026
