Totoro: A Scalable Federated Learning Engine for the Edge
Cheng-Wei Ching, Xin Chen, Taehwan Kim, Bo Ji, Qingyang Wang, Dilma Da Silva, Liting Hu
摘要
Federated Learning (FL) is an emerging distributed machine learning (ML) technique that enables in-situ model training and inference on decentralized edge devices. We propose Totoro, a novel scalable FL engine, that enables massive FL applications to run simultaneously on edge networks. The key insight is to explore a distributed hash table (DHT)-based peer-to-peer (P2P) model to re-architect the centralized FL system design into a fully decentralized one. In contrast to previous studies where many FL applications shared one centralized parameter server, Totoro assigns a dedicated parameter server to each individual application. Any edge node can act as any application's coordinator, aggregator, client selector, worker (participant device), or any combination of the above, thereby radically improving scalability and adaptivity. Totoro introduces three innovations to realize its design: a locality-aware P2P multi-ring structure, a publish/subscribe-based forest abstraction, and a bandit-based exploitation-exploration path planning model. Real-world experiments on 500 Amazon EC2 servers show that Totoro scales gracefully with the number of FL applications and N edge nodes, speeds up the total training time by 1.2 × -14.0×, achieves O (logN) hops for model dissemination and gradient aggregation with millions of nodes, and efficiently adapts to the practical edge networks and churns.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper17
- Adaptive Federated OptimizationSashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett 等ICLR 2021 · 被引用 1,917 次
- Fair Resource Allocation in Federated LearningTian Li, Maziar Sanjabi, Ahmad Beirami, Virginia SmithICLR 2020 · 被引用 971 次
- BatchCrypt: Efficient Homomorphic Encryption for Cross-Silo Federated LearningChengliang Zhang, Suyi Li, Junzhe Xia, Wei Wang 等USENIX ATC 2020 · 被引用 967 次
- Don't Use Large Mini-batches, Use Local SGDTao Lin, Sebastian U. Stich, Kumar Kshitij Patel, Martin JaggiICLR 2020 · 被引用 462 次
- FetchSGD: Communication-Efficient Federated Learning with SketchingDaniel Rothchild, Ashwinee Panda, Enayat Ullah, Nikita Ivkin 等ICML 2020 · 被引用 425 次
相关 Paper
- Adaptive Configuration for Heterogeneous Participants in Decentralized Federated LearningYunming Liao, Yang Xu, Hongli Xu, Lun Wang 等INFOCOM 2023 · 被引用 66 次
- MAS: Towards Resource-Efficient Federated Multiple-Task LearningWeiming Zhuang, Yonggang Wen, Lingjuan Lyu, Shuai ZhangICCV 2023 · 被引用 22 次
- Oort: Efficient Federated Learning via Guided Participant SelectionFan Lai, Xiangfeng Zhu, Harsha V. Madhyastha, Mosharaf ChowdhuryOSDI 2021
- Enhancing Decentralized Federated Learning for Non-IID Data on Heterogeneous DevicesMin Chen, Yang Xu, Hongli Xu, Liusheng HuangICDE 2023 · 被引用 25 次
- FlexiFed: Personalized Federated Learning for Edge Clients with Heterogeneous Model ArchitecturesKaibin Wang, Qiang He, Feifei Chen, Chunyang Chen 等WWW 2023 · 被引用 69 次
