AVM-BTB: Adaptive and Virtualized Multi-level Branch Target Buffer
Yunzhe Liu, Xinyu Li, Tingting Zhang, Tianyi Liu, Qi Guo, Fuxin Zhang, Jian Wang
Abstract
Branch Target Buffer (BTB) plays an important role in modern processors. It is used to identify branches in the instruction stream and predict branch targets. The accuracy of BTB is highly impacted by BTB capacity. However, expanding BTB capacity using traditional methods requires valuable on-chip SRAM. Both timing and area restriction make these approaches unsustainable. Moreover, these methods overlook the different demands of various applications, leading to increased power consumption and resource waste in some cases. To address this problem, we propose AVM-BTB. The key observations behind AVM-BTB come from three aspects: 1) BTB requirements vary over different applications and even over different running stages of the same application. 2) Micro-operation Cache (Uop Cache) and ICache exhibit inefficiency when confronted with instruction footprints that greatly exceed their capacity. 3) In specific scenarios of frontend overload, reducing cache capacity and increasing BTB size can effectively mitigate expensive branch prediction errors. Simultaneously, the implementation of Fetch Directed Instruction Prefetching (FDIP) can offset the limitations in cache capacity to some extent. These observations reveal the feasibility of dynamically borrowing cache capacity as temporary BTB and returning these BTB to cache when they are not needed, further resulting in an adaptive and virtualized multi-level BTB scheme. However, such a BTB structure is non-trivial. In this work, from the perspective of instructions, the cache hierarchy stores instruction data, while the BTB stores metadata used for branch prediction and instruction prefetch. Targeting high performance, AVM-BTB maintains a dynamic balance between data and metadata by monitoring the BTB error rate and effective accesses. Evaluation with 1253 traces shows that AVM-BTB is suitable for both frontend-bound and frontend-friendly scenarios, without consuming additional SRAM and with reasonable implementation efforts. Compared to baseline, AVM-BTB delivers an average performance boost of and a power consumption reduction of . It also outperforms the five state-of-the-art solutions by on average in terms of IPC.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get efae20c0-01d9-4390-8a87-c7f266c4c430Cited by top-tier papers2
- Assassyn: A Unified Abstraction for Architectural Simulation and ImplementationJian Weng, Boyang Han, Derui Gao, Ruijie Gao et al.ISCA 2025 · 1 citation
- CryptoBTB: A Secure Hierarchical BTB for Diverse Instruction Footprint WorkloadsDebpratim Adak, Eric Rotenberg, Amro Awad, Huiyang ZhouMICRO 2025 · 1 citation
Related papers
- Skia: Exposing Shadow BranchesChrysanthos Pepi, Bhargav Reddy Godala, Krishnam Tibrewala, Gino A. Chacon et al.ASPLOS 2025 · 2 citations
- A Storage-Effective BTB Organization for ServersTruls Asheim, Boris Grot, Rakesh KumarHPCA 2023 · 12 citations
- Branch Target Buffer OrganizationsArthur Perais, Rami SheikhMICRO 2023 · 8 citations
- Twig: Profile-Guided BTB Prefetching for Data Center ApplicationsTanvir Ahmed Khan, Nathan Brown, Akshitha Sriraman, Niranjan K. Soundararajan et al.MICRO 2021 · 33 citations
- PDede: Partitioned, Deduplicated, Delta Branch Target BufferNiranjan K. Soundararajan, Peter Braun, Tanvir Ahmed Khan, Baris Kasikci et al.MICRO 2021 · 22 citations
