Auto-Predication of Critical Branches
Adarsh Chauhan, Jayesh Gaur, Zeev Sperber, Franck Sala, Lihu Rappoport, Adi Yoaz, Sreenivas Subramoney
摘要
Advancements in branch predictors have allowed modern processors to aggressively speculate and gain significant performance with every generation of increasing out-of-order depth and width. Unfortunately, there are branches that are still hard-to-predict (H2P) and mis-speculation on these branches is severely limiting the performance scalability of future processors. One potential solution to mitigate this problem is to predicate branches by substituting control dependencies with data dependencies. Predication is very costly for performance as it inhibits instruction level parallelism. To overcome this limitation, prior works selectively applied predication at run-time on H2P branches that have low confidence of branch prediction. However, these schemes do not fully comprehend the delicate trade-offs involved in suppressing speculation and can suffer from performance degradation on certain workloads. Additionally, they need significant changes not just to the hardware but also to the compiler and the instruction set architecture, rendering their implementation complex and challenging.In this paper, by analyzing the fundamental trade-offs between branch prediction and predication, we propose Auto-Predication of Critical Branches (ACB) - an end-to-end hardware-based solution that intelligently disables speculation only on branches that are critical for performance. Unlike existing approaches, ACB uses a sophisticated performance monitoring mechanism to gauge the effectiveness of dynamic predication, and hence does not suffer from performance inversions. Our simulation results show that, with just 386 bytes of additional hardware and no software support, ACB delivers 8% performance gain over a baseline similar to the Skylake processor. We also show that ACB reduces pipeline flushes because of mis-speculations by 22%, thus effectively helping both power and performance.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Branch Runahead: An Alternative to Branch Prediction for Impossible to Predict BranchesStephen Pruett, Yale N. PattMICRO 2021 · 被引用 23 次
- Enabling Branch-Mispredict Level Parallelism by Selectively Flushing InstructionsStijn Eyerman, Wim Heirman, Sam Van den Steen, Ibrahim HurMICRO 2021 · 被引用 10 次
- Alternate Path μ-op Cache PrefetchingSawan Singh, Arthur Perais, Alexandra Jimborean, Alberto RosISCA 2024 · 被引用 4 次
- Alternate Path FetchAniket Deshmukh, Lingzhe Chester Cai, Yale N. PattISCA 2024 · 被引用 2 次
相关 Paper
- Timely, Efficient, and Accurate Branch PrecomputationAniket Deshmukh, Lingzhe Chester Cai, Yale N. PattMICRO 2024 · 被引用 3 次
- Focused Value PredictionSumeet Bandishte, Jayesh Gaur, Zeev Sperber, Lihu Rappoport 等ISCA 2020 · 被引用 13 次
- CRISP: critical slice prefetchingHeiner Litz, Grant Ayers, Parthasarathy RanganathanASPLOS 2022 · 被引用 33 次
- The Last-Level Branch PredictorDavid Schall, Andreas Sandberg, Boris GrotMICRO 2024 · 被引用 10 次
- Boosting Store Buffer Efficiency with Store-Prefetch BurstsJuan M. Cebrian, Stefanos Kaxiras, Alberto RosMICRO 2020 · 被引用 4 次
