Divide and Conquer Frontend Bottleneck
Ali Ansari, Pejman Lotfi-Kamran, Hamid Sarbazi-Azad
摘要
The frontend stalls caused by instruction and BTB misses are a significant source of performance degradation in server processors. Prefetchers are commonly employed to mitigate frontend bottleneck. However, next-line prefetchers, which are available in server processors, are incapable of eliminating a considerable number of L1 instruction misses. Temporal instruction prefetchers, on the other hand, effectively remove most of the instruction and BTB misses but impose significant area overhead.Recently, an old idea of using BTB-directed instruction prefetching is revived to address the limitations of temporal instruction prefetchers. While this approach leads to prefetchers with low area overhead, it requires significant changes to the frontend of a processor. Moreover, as this approach relies on the BTB content for prefetching, BTB misses stall the prefetcher, and likely lead to costly instruction misses. Especially as instruction misses are usually more expensive than BTB misses, the dependence of instruction prefetching to the BTB content is harmful to workloads with very large instruction footprints. Moreover, BTB-directed instruction prefetchers, as proposed in prior work, cannot be applied to variable-length ISAs.In this work, we showcase the harmful effects of making instruction prefetchers depend on the BTB content. Moreover, we divide the frontend bottleneck into three categories and use a divide-and-conquer approach to propose simple and effective solutions for each one. Sequential misses can be covered by an accurate and timely sequential prefetcher named SN4L, a lightweight discontinuity prefetcher named Dis eliminates discontinuity misses, and the BTB misses are reduced by pre-decoding the prefetched blocks. We also discuss how our proposal can be used for variable-length ISAs with low storage overhead. Our proposal, SN4L+ Dis+BTB, imposes the same area overhead as the state-of-the-art BTB-directed prefetcher, and at the same time, outperforms it by 5% on average and up to 16%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Lukewarm serverless functions: characterization and optimizationDavid Schall, Artemiy Margaritov, Dmitrii Ustiugov, Andreas Sandberg 等ISCA 2022 · 被引用 36 次
- Twig: Profile-Guided BTB Prefetching for Data Center ApplicationsTanvir Ahmed Khan, Nathan Brown, Akshitha Sriraman, Niranjan K. Soundararajan 等MICRO 2021 · 被引用 33 次
- Cerebros: Evading the RPC Tax in DatacentersArash Pourhabibi Zarandi, Mark Sutherland, Alexandros Daglis, Babak FalsafiMICRO 2021 · 被引用 24 次
- Thermometer: profile-guided btb replacement for data center applicationsShixin Song, Tanvir Ahmed Khan, Sara Mahdizadeh-Shahri, Akshitha Sriraman 等ISCA 2022 · 被引用 23 次
- PDede: Partitioned, Deduplicated, Delta Branch Target BufferNiranjan K. Soundararajan, Peter Braun, Tanvir Ahmed Khan, Baris Kasikci 等MICRO 2021 · 被引用 22 次
相关 Paper
- Morrigan: A Composite Instruction TLB PrefetcherGeorgios Vavouliotis, Lluc Alvarez, Boris Grot, Daniel A. Jiménez 等MICRO 2021 · 被引用 19 次
- Skia: Exposing Shadow BranchesChrysanthos Pepi, Bhargav Reddy Godala, Krishnam Tibrewala, Gino A. Chacon 等ASPLOS 2025 · 被引用 2 次
- PDIP: Priority Directed Instruction PrefetchingBhargav Reddy Godala, Sankara Prasad Ramesh, Gilles A. Pokam, Jared Stark 等ASPLOS 2024 · 被引用 17 次
- Branch Target Buffer OrganizationsArthur Perais, Rami SheikhMICRO 2023 · 被引用 8 次
- Enhancing Instruction Prefetching via Cache and TLB ManagementAlexandre Valentin Jamet, Georgios Vavouliotis, Martí Torrents, Dimitrios Chasapis 等ISCA 2026
