Owl: Scale and Flexibility in Distribution of Hot Content
Jason Flinn, Xianzheng Dou, Arushi Aggarwal, Alex Boyko, Francois Richard, Eric Sun, Wendy Tobagus, Nick Wolchko, Fang Zhou
摘要
Owl provides high-fanout distribution of large data objects to hosts in Meta's private cloud. Owl combines a decentralized data plane based on ephemeral peer-to-peer distribution trees with a centralized control plane in which tracker services maintain detailed metadata about peers, their cache state, and ongoing downloads. In Owl, peer nodes are simple state machines and centralized trackers decide from where each peer should fetch data, how they should retry on failure, and which data they should cache and evict. Owl trackers provide a highly-flexible and configurable policy interface that customizes and optimizes behavior for widely-varying distribution use cases. In contrast to prior assumptions about peer-to-peer distribution, Owl shows that centralizing the control plan is not a barrier to scalability: Owl distributes over 800 petabytes of data per day to millions of client processes. Owl improves download speeds by a factor of 2-3 over both BitTorrent and a prior decentralized static distribution tree used at Meta, while supporting 106 use cases that collectively employ 55 different distribution policies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- SimAI: Unifying Architecture Design and Performance Tuning for Large-Scale Large Language Model Training with Scalability and PrecisionXizheng Wang, Qingxu Li, Yichi Xu, Gang Lu 等NSDI 2025 · 被引用 82 次
- XFaaS: Hyperscale and Low Cost Serverless Functions at MetaAlireza Sahraei, Soteris Demetriou, Amirali Sobhgol, Haoran Zhang 等SOSP 2023 · 被引用 32 次
- Remote Procedure Call as a Managed System ServiceJingrong Chen, Yongji Wu, Shihan Lin, Yechen Xu 等NSDI 2023 · 被引用 30 次
- Cloudcast: High-Throughput, Cost-Aware Overlay Multicast in the CloudSarah Wooders, Shu Liu, Paras Jain, Xiangxi Mo 等NSDI 2024 · 被引用 20 次
- RobustRL: Role-Based Fault Tolerance System for RL Post-TrainingZhenqian Chen, Baoquan Zhong, Xiang Li, Qing Dai 等OSDI 2026
它引用的顶会 Paper6
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 被引用 245 次
- FaaSNet: Scalable and Fast Provisioning of Custom Serverless Container Runtimes at Alibaba Cloud Function ComputeAo Wang, Shuai Chang, Huangshi Tian, Hongqi Wang 等USENIX ATC 2021 · 被引用 171 次
- The CacheLib Caching Engine: Design and Experiences at ScaleBenjamin Berg, Daniel S. Berger, Sara McAllister, Isaac Grosof 等OSDI 2020 · 被引用 145 次
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor 等OSDI 2020 · 被引用 107 次
- Orion: Google's Software-Defined Networking Control PlaneAndrew D. Ferguson, Steve D. Gribble, Chi-Yao Hong, Charles Killian 等NSDI 2021 · 被引用 95 次
相关 Paper
- Accelerating Metadata Management of DFS via Speculative Permission CheckingYiduo Wang, Linghang Meng, Liang Li, Jie WuICDE 2026
- An exabyte a day: throughput-oriented, large scale, managed data transfers with EffingoLadislav Pápay, Jan Pustelnik, Krzysztof Rzadca, Beata Strack 等SIGCOMM 2024 · 被引用 11 次
- EBB: Reliable and Evolvable Express Backbone Network in MetaMarek Denis, Yuanjun Yao, Ashley Hatch, Qin Zhang 等SIGCOMM 2023 · 被引用 25 次
- Design and evaluation of IPFS: a storage layer for the decentralized webDennis Trautwein, Aravindh Raman, Gareth Tyson, Ignacio Castro 等SIGCOMM 2022 · 被引用 188 次
- Network entitlement: contract-based network sharing with agility and SLO guaranteesSatyajeet Singh Ahuja, Vinayak Dangui, Kirtesh Patil, Manikandan Somasundaram 等SIGCOMM 2022 · 被引用 3 次
