Owl: Scale and Flexibility in Distribution of Hot Content
Jason Flinn, Xianzheng Dou, Arushi Aggarwal, Alex Boyko, Francois Richard, Eric Sun, Wendy Tobagus, Nick Wolchko, Fang Zhou
Abstract
Owl provides high-fanout distribution of large data objects to hosts in Meta's private cloud. Owl combines a decentralized data plane based on ephemeral peer-to-peer distribution trees with a centralized control plane in which tracker services maintain detailed metadata about peers, their cache state, and ongoing downloads. In Owl, peer nodes are simple state machines and centralized trackers decide from where each peer should fetch data, how they should retry on failure, and which data they should cache and evict. Owl trackers provide a highly-flexible and configurable policy interface that customizes and optimizes behavior for widely-varying distribution use cases. In contrast to prior assumptions about peer-to-peer distribution, Owl shows that centralizing the control plan is not a barrier to scalability: Owl distributes over 800 petabytes of data per day to millions of client processes. Owl improves download speeds by a factor of 2-3 over both BitTorrent and a prior decentralized static distribution tree used at Meta, while supporting 106 use cases that collectively employ 55 different distribution policies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- SimAI: Unifying Architecture Design and Performance Tuning for Large-Scale Large Language Model Training with Scalability and PrecisionXizheng Wang, Qingxu Li, Yichi Xu, Gang Lu et al.NSDI 2025 · 82 citations
- XFaaS: Hyperscale and Low Cost Serverless Functions at MetaAlireza Sahraei, Soteris Demetriou, Amirali Sobhgol, Haoran Zhang et al.SOSP 2023 · 32 citations
- Remote Procedure Call as a Managed System ServiceJingrong Chen, Yongji Wu, Shihan Lin, Yechen Xu et al.NSDI 2023 · 30 citations
- Cloudcast: High-Throughput, Cost-Aware Overlay Multicast in the CloudSarah Wooders, Shu Liu, Paras Jain, Xiangxi Mo et al.NSDI 2024 · 20 citations
- RobustRL: Role-Based Fault Tolerance System for RL Post-TrainingZhenqian Chen, Baoquan Zhong, Xiang Li, Qing Dai et al.OSDI 2026
Builds on6
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 245 citations
- FaaSNet: Scalable and Fast Provisioning of Custom Serverless Container Runtimes at Alibaba Cloud Function ComputeAo Wang, Shuai Chang, Huangshi Tian, Hongqi Wang et al.USENIX ATC 2021 · 171 citations
- The CacheLib Caching Engine: Design and Experiences at ScaleBenjamin Berg, Daniel S. Berger, Sara McAllister, Isaac Grosof et al.OSDI 2020 · 145 citations
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor et al.OSDI 2020 · 107 citations
- Orion: Google's Software-Defined Networking Control PlaneAndrew D. Ferguson, Steve D. Gribble, Chi-Yao Hong, Charles Killian et al.NSDI 2021 · 95 citations
Related papers
- Accelerating Metadata Management of DFS via Speculative Permission CheckingYiduo Wang, Linghang Meng, Liang Li, Jie WuICDE 2026
- An exabyte a day: throughput-oriented, large scale, managed data transfers with EffingoLadislav Pápay, Jan Pustelnik, Krzysztof Rzadca, Beata Strack et al.SIGCOMM 2024 · 11 citations
- EBB: Reliable and Evolvable Express Backbone Network in MetaMarek Denis, Yuanjun Yao, Ashley Hatch, Qin Zhang et al.SIGCOMM 2023 · 25 citations
- Design and evaluation of IPFS: a storage layer for the decentralized webDennis Trautwein, Aravindh Raman, Gareth Tyson, Ignacio Castro et al.SIGCOMM 2022 · 188 citations
- Network entitlement: contract-based network sharing with agility and SLO guaranteesSatyajeet Singh Ahuja, Vinayak Dangui, Kirtesh Patil, Manikandan Somasundaram et al.SIGCOMM 2022 · 3 citations
