Sol: Fast Distributed Computation Over Slow Networks
Fan Lai, Jie You, Xiangfeng Zhu, Harsha V. Madhyastha, Mosharaf Chowdhury
摘要
The popularity of big data and AI has led to many optimizations at different layers of distributed computation stacks. Despite -or perhaps, because of -its role as the narrow waist of such software stacks, the design of the execution engine, which is in charge of executing every single task of a job, has mostly remained unchanged. As a result, the execution engines available today are ones primarily designed for low latency and high bandwidth datacenter networks. When either or both of the network assumptions do not hold, CPUs are significantly underutilized.
In this paper, we take a first-principles approach toward developing an execution engine that can adapt to diverse network conditions. Sol, our federated execution engine architecture, flips the status quo in two respects. First, to mitigate the impact of high latency, Sol proactively assigns tasks, but does so judiciously to be resilient to uncertainties. Second, to improve the overall resource utilization, Sol decouples communication from computation internally instead of committing resources to both aspects of a task simultaneously. Our evaluations on EC2 show that, compared to Apache Spark in resource-constrained networks, Sol improves SQL and machine learning jobs by 16.4× and 4.2× on average.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- ModelKeeper: Accelerating DNN Training via Automated Training WarmupFan Lai, Yinwei Dai, Harsha V. Madhyastha, Mosharaf ChowdhuryNSDI 2023 · 被引用 31 次
- GREEN: Carbon-efficient Resource Scheduling for Machine Learning ClustersKaiqiang Xu, Decang Sun, Han Tian, Junxue Zhang 等NSDI 2025 · 被引用 23 次
- Totoro: A Scalable Federated Learning Engine for the EdgeCheng-Wei Ching, Xin Chen, Taehwan Kim, Bo Ji 等EuroSys 2024 · 被引用 12 次
- Saba: Rethinking Datacenter Network Allocation from Application's PerspectiveM. R. Siavash Katebzadeh, Paolo Costa, Boris GrotEuroSys 2023 · 被引用 4 次
- Sequoia: An Accessible and Extensible Framework for Privacy-Preserving Machine Learning over Distributed DataKaiqiang Xu, Di Chai, Junxue Zhang, Fan Lai 等SIGMOD 2025 · 被引用 1 次
它引用的顶会 Paper1
相关 Paper
- Accio: Bolt-on Query FederationXiaoying Wang, Jiannan Wang, Tianzheng Wang, Yong ZhangVLDB 2025
- Generalized Sub-Query Fusion for Eliminating Redundant I/O from Big-Data QueriesPartho Sarthi, Kaushik Rajan, Akash Lal, Abhishek Modi 等OSDI 2020 · 被引用 5 次
- Decouple and Decompose: Scaling Resource Allocation with DeDeZhiying Xu, Minlan Yu, Francis Y. YanOSDI 2025 · 被引用 5 次
- Breaking the computation and communication abstraction barrier in distributed machine learning workloadsAbhinav Jangda, Jun Huang, Guodong Liu, Amir Hossein Nodehi Sabet 等ASPLOS 2022 · 被引用 68 次
- Intra-Query Runtime Elasticity for Cloud-Native Data AnalysisXukang Zhang, Huanchen Zhang, Xiaofeng MengSIGMOD 2025 · 被引用 3 次
