A Storage-Effective BTB Organization for Servers
Truls Asheim, Boris Grot, Rakesh Kumar
Abstract
Many contemporary applications feature multimegabyte instruction footprints that overwhelm the capacity of branch target buffers (BTB) and instruction caches (L1-I), causing frequent front-end stalls that inevitably hurt performance. BTB capacity is crucial for performance as a sufficiently large BTB enables the front-end to accurately resolve the upcoming execution path and steer instruction fetch appropriately. Moreover, it also enables highly effective fetch-directed instruction prefetching that can eliminate a large portion L1-I misses. For these reasons, commercial processors allocate vast amounts of storage capacity to BTBs.
This work aims to reduce BTB storage requirements by optimizing the organization of BTB entries. Our key insight is that storing branch target offsets, instead of full or compressed targets, can drastically reduce BTB storage cost as the vast majority of dynamic branches have short offsets requiring just a handful of bits to encode. Based on this insight, we size the ways of a set associative BTB to hold different number of target offset bits such that each way stores offsets within a particular range. Doing so enables a dramatic reduction in storage for target addresses. Our final design, called BTB-X, uses an 8-way set associative BTB with differently sized ways that enables it to track about 2.24x more branches than a conventional BTB and 1.3x more branches than a storage-optimized state-of-the-art BTB organization, called PDede, with the same storage budget.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9810c441-000c-41a8-a60b-52a15b98d44cCited by top-tier papers6
- Phantom: Exploiting Decoder-detectable MispredictionsJohannes Wikner, Daniël Trujillo, Kaveh RazaviMICRO 2023 · 19 citations
- Warming Up a Cold Front-End with IgniteDavid Schall, Andreas Sandberg, Boris GrotMICRO 2023 · 11 citations
- Branch Target Buffer OrganizationsArthur Perais, Rami SheikhMICRO 2023 · 8 citations
- Alternate Path μ-op Cache PrefetchingSawan Singh, Arthur Perais, Alexandra Jimborean, Alberto RosISCA 2024 · 4 citations
- Weeding out Front-End Stalls with Uneven Block Size Instruction CacheRoman Brunner, Rakesh KumarMICRO 2024 · 1 citation
Builds on4
- I-SPY: Context-Driven Conditional Instruction Prefetching with CoalescingTanvir Ahmed Khan, Akshitha Sriraman, Joseph Devietti, Gilles Pokam et al.MICRO 2020 · 37 citations
- Lukewarm serverless functions: characterization and optimizationDavid Schall, Artemiy Margaritov, Dmitrii Ustiugov, Andreas Sandberg et al.ISCA 2022 · 36 citations
- Twig: Profile-Guided BTB Prefetching for Data Center ApplicationsTanvir Ahmed Khan, Nathan Brown, Akshitha Sriraman, Niranjan K. Soundararajan et al.MICRO 2021 · 33 citations
- PDede: Partitioned, Deduplicated, Delta Branch Target BufferNiranjan K. Soundararajan, Peter Braun, Tanvir Ahmed Khan, Baris Kasikci et al.MICRO 2021 · 22 citations
Related papers
- AVM-BTB: Adaptive and Virtualized Multi-level Branch Target BufferYunzhe Liu, Xinyu Li, Tingting Zhang, Tianyi Liu et al.ISCA 2024 · 6 citations
- Skia: Exposing Shadow BranchesChrysanthos Pepi, Bhargav Reddy Godala, Krishnam Tibrewala, Gino A. Chacon et al.ASPLOS 2025 · 2 citations
- Divide and Conquer Frontend BottleneckAli Ansari, Pejman Lotfi-Kamran, Hamid Sarbazi-AzadISCA 2020 · 28 citations
- Thermometer: profile-guided btb replacement for data center applicationsShixin Song, Tanvir Ahmed Khan, Sara Mahdizadeh-Shahri, Akshitha Sriraman et al.ISCA 2022 · 23 citations
- PDIP: Priority Directed Instruction PrefetchingBhargav Reddy Godala, Sankara Prasad Ramesh, Gilles A. Pokam, Jared Stark et al.ASPLOS 2024 · 17 citations
