Weeding out Front-End Stalls with Uneven Block Size Instruction Cache
Roman Brunner, Rakesh Kumar
摘要
The core front-end remains a critical bottleneck in modern server workloads owing to their multi-MB instruction footprints stemming from deep software stacks. Prior work has mainly investigated instruction prefetching and cache replacement policies to mitigate this bottleneck. In this work, we take an orthogonal approach and analyze instruction cache storage efficiency. Our analysis shows that, on average, about 60% of the bytes in a cache block are never accessed before the block is evicted from the instruction cache. This represents a huge storage inefficiency that more than halves the effective cache capacity. We observe that this inefficiency is caused by the fixed cache block sizes which are unable to accommodate the varying spatial locality inherent in the instruction stream. To mitigate this inefficiency, we propose Uneven Block Size (UBS) instruction cache, which supports different cache block sizes in a cache set. Our evaluation shows that UBS cache improves the storage efficiency by 32 percentage points over the baseline instruction cache. Further, by supporting uneven block sizes, UBS cache accommodates more than twice the number of blocks than a conventional cache within a given storage budget. Overall, the additional blocks combined with the better storage efficiency result in UBS cache approaching the performance of a 64KB conventional cache on a set of server workloads while requiring a storage budget similar to a 32KB conventional cache.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- I-SPY: Context-Driven Conditional Instruction Prefetching with CoalescingTanvir Ahmed Khan, Akshitha Sriraman, Joseph Devietti, Gilles Pokam 等MICRO 2020 · 被引用 37 次
- Lukewarm serverless functions: characterization and optimizationDavid Schall, Artemiy Margaritov, Dmitrii Ustiugov, Andreas Sandberg 等ISCA 2022 · 被引用 36 次
- Ripple: Profile-Guided Instruction Cache Replacement for Data Center ApplicationsTanvir Ahmed Khan, Dexin Zhang, Akshitha Sriraman, Joseph Devietti 等ISCA 2021 · 被引用 33 次
- Twig: Profile-Guided BTB Prefetching for Data Center ApplicationsTanvir Ahmed Khan, Nathan Brown, Akshitha Sriraman, Niranjan K. Soundararajan 等MICRO 2021 · 被引用 33 次
- A Cost-Effective Entangling Prefetcher for InstructionsAlberto Ros, Alexandra JimboreanISCA 2021 · 被引用 31 次
相关 Paper
- Divide and Conquer Frontend BottleneckAli Ansari, Pejman Lotfi-Kamran, Hamid Sarbazi-AzadISCA 2020 · 被引用 28 次
- Garibaldi: A Pairwise Instruction-Data Management for Enhancing Shared Last-Level Cache Performance in Server WorkloadsJaewon Kwon, Yongju Lee, Jiwan Kim, Enhyeok Jang 等ISCA 2025 · 被引用 1 次
- PDede: Partitioned, Deduplicated, Delta Branch Target BufferNiranjan K. Soundararajan, Peter Braun, Tanvir Ahmed Khan, Baris Kasikci 等MICRO 2021 · 被引用 22 次
- Enhancing Instruction Prefetching via Cache and TLB ManagementAlexandre Valentin Jamet, Georgios Vavouliotis, Martí Torrents, Dimitrios Chasapis 等ISCA 2026
- PDIP: Priority Directed Instruction PrefetchingBhargav Reddy Godala, Sankara Prasad Ramesh, Gilles A. Pokam, Jared Stark 等ASPLOS 2024 · 被引用 17 次
