Propeller: A Profile Guided, Relinking Optimizer for Warehouse-Scale Applications
Han Shen, Krzysztof Pszeniczny, Rahman Lavaee, Snehasish Kumar, Sriraman Tallam, Xinliang David Li
摘要
While profile guided optimizations (PGO) and link time optimiza-tions (LTO) have been widely adopted, post link optimizations (PLO)have languished until recently when researchers demonstrated that late injection of profiles can yield significant performance improvements. However, the disassembly-driven, monolithic design of post link optimizers face scaling challenges with large binaries andis at odds with distributed build systems. To reconcile and enable post link optimizations within a distributed build environment, we propose Propeller, a relinking optimizer for warehouse scale work-loads. To enable flexible code layout optimizations, we introduce basic block sections, a novel linker abstraction. Propeller uses basic block sections to enable a new approach to PLO without disassembly. Propeller achieves scalability by relinking the binary using precise profiles instead of rewriting the binary. The overhead of relinking is lowered by caching and leveraging distributed compiler actions during code generation. Propeller has been deployed to production at Google with over tens of millions of cores executing Propeller optimized code at any time. An evaluation of internal warehouse-scale applications show Propeller improves performance by 1.1% to 8% beyond PGO and ThinLTO. Compiler tools such as Clang improve by 7% while MySQL improves by 1%. Compared to the state of the art binary optimizer, Propeller achieves comparable performance while lowering memory overheads by 30%-70% on large benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Tiered Memory Management Beyond HotnessJinshu Liu, Hamid Hadian, Hanchen Xu, Huaicheng LiOSDI 2025 · 被引用 13 次
- Getting a Handle on Unmanaged MemoryNick Wanninger, Tommy McMichen, Simone Campanoni, Peter A. DindaASPLOS 2024 · 被引用 3 次
- Wax: Optimizing Data Center Applications With Stale ProfileTawhid Bhuiyan, Sumya Hoque, Angelica Aparecida Moreira, Tanvir Ahmed KhanASPLOS 2026 · 被引用 1 次
- Diatom: Polylithic Binary Lifting with Data-Flow Summaries and Type-Aware IR LinkingAnshunkang Zhou, Charles ZhangOOPSLA 2026 · 被引用 1 次
- A TRRIP Down Memory Lane: Temperature-Based Re-Reference Interval Prediction For Instruction CachingHenry Kao, Nikhil Sreekumar, Prabhdeep Singh Soni, Ali Sedaghati 等MICRO 2025 · 被引用 1 次
它引用的顶会 Paper3
- An In-Depth Analysis of Disassembly on Full-Scale x86/x64 BinariesDennis Andriesse, Xi Chen, Victor van der Veen, Asia Slowinska 等USENIX Security 2016 · 被引用 162 次
- Beyond malloc efficiency to fleet efficiency: a hugepage-aware memory allocatorA. H. Hunter, Chris Kennelly, Paul Turner, Darryl Gove 等OSDI 2021 · 被引用 51 次
- Scalable build service system with smart scheduling serviceKaiyuan Wang, Greg Tener, Vijay Gullapalli, Xin Huang 等ISSTA 2020 · 被引用 8 次
相关 Paper
- Profile inference revisitedWenlei He, Julián Mestre, Sergey Pupyrev, Lei Wang 等POPL 2022 · 被引用 16 次
- VESPA: static profiling for binary optimizationAngelica Aparecida Moreira, Guilherme Ottoni, Fernando Magno Quintão PereiraOOPSLA 2021 · 被引用 18 次
- Automatic Propagation of Profile Information through the Optimization PipelineElisa Fröhlich, Angelica Aparecida Moreira, Fernando Magno Quintão PereiraOOPSLA 2026
- ProfiX: Improving Profile-Guided Optimization in Compilers with Graph Neural NetworksHuiri Tan, Juyong Jiang, Jiasi ShenNeurIPS 2025 · 被引用 4 次
- TypeCraft: A Lightweight Data Type Profiler with High ResolutionZecheng Li, Xu Liu, Namhyung Kim, Blake Jones 等OSDI 2026
