Propeller: A Profile Guided, Relinking Optimizer for Warehouse-Scale Applications
Han Shen, Krzysztof Pszeniczny, Rahman Lavaee, Snehasish Kumar, Sriraman Tallam, Xinliang David Li
Abstract
While profile guided optimizations (PGO) and link time optimiza-tions (LTO) have been widely adopted, post link optimizations (PLO)have languished until recently when researchers demonstrated that late injection of profiles can yield significant performance improvements. However, the disassembly-driven, monolithic design of post link optimizers face scaling challenges with large binaries andis at odds with distributed build systems. To reconcile and enable post link optimizations within a distributed build environment, we propose Propeller, a relinking optimizer for warehouse scale work-loads. To enable flexible code layout optimizations, we introduce basic block sections, a novel linker abstraction. Propeller uses basic block sections to enable a new approach to PLO without disassembly. Propeller achieves scalability by relinking the binary using precise profiles instead of rewriting the binary. The overhead of relinking is lowered by caching and leveraging distributed compiler actions during code generation. Propeller has been deployed to production at Google with over tens of millions of cores executing Propeller optimized code at any time. An evaluation of internal warehouse-scale applications show Propeller improves performance by 1.1% to 8% beyond PGO and ThinLTO. Compiler tools such as Clang improve by 7% while MySQL improves by 1%. Compared to the state of the art binary optimizer, Propeller achieves comparable performance while lowering memory overheads by 30%-70% on large benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e08e8deb-4ede-41e9-9b0f-6686ce0cf2acCited by top-tier papers9
- Tiered Memory Management Beyond HotnessJinshu Liu, Hamid Hadian, Hanchen Xu, Huaicheng LiOSDI 2025 · 13 citations
- Getting a Handle on Unmanaged MemoryNick Wanninger, Tommy McMichen, Simone Campanoni, Peter A. DindaASPLOS 2024 · 3 citations
- Wax: Optimizing Data Center Applications With Stale ProfileTawhid Bhuiyan, Sumya Hoque, Angelica Aparecida Moreira, Tanvir Ahmed KhanASPLOS 2026 · 1 citation
- Diatom: Polylithic Binary Lifting with Data-Flow Summaries and Type-Aware IR LinkingAnshunkang Zhou, Charles ZhangOOPSLA 2026 · 1 citation
- A TRRIP Down Memory Lane: Temperature-Based Re-Reference Interval Prediction For Instruction CachingHenry Kao, Nikhil Sreekumar, Prabhdeep Singh Soni, Ali Sedaghati et al.MICRO 2025 · 1 citation
Builds on3
- An In-Depth Analysis of Disassembly on Full-Scale x86/x64 BinariesDennis Andriesse, Xi Chen, Victor van der Veen, Asia Slowinska et al.USENIX Security 2016 · 162 citations
- Beyond malloc efficiency to fleet efficiency: a hugepage-aware memory allocatorA. H. Hunter, Chris Kennelly, Paul Turner, Darryl Gove et al.OSDI 2021 · 51 citations
- Scalable build service system with smart scheduling serviceKaiyuan Wang, Greg Tener, Vijay Gullapalli, Xin Huang et al.ISSTA 2020 · 8 citations
Related papers
- Profile inference revisitedWenlei He, Julián Mestre, Sergey Pupyrev, Lei Wang et al.POPL 2022 · 16 citations
- VESPA: static profiling for binary optimizationAngelica Aparecida Moreira, Guilherme Ottoni, Fernando Magno Quintão PereiraOOPSLA 2021 · 18 citations
- Automatic Propagation of Profile Information through the Optimization PipelineElisa Fröhlich, Angelica Aparecida Moreira, Fernando Magno Quintão PereiraOOPSLA 2026
- ProfiX: Improving Profile-Guided Optimization in Compilers with Graph Neural NetworksHuiri Tan, Juyong Jiang, Jiasi ShenNeurIPS 2025 · 4 citations
- TypeCraft: A Lightweight Data Type Profiler with High ResolutionZecheng Li, Xu Liu, Namhyung Kim, Blake Jones et al.OSDI 2026
