Tiny Tasks - A Remedy for Synchronization Constraints in Multi-Server Systems
Markus Fidler, Brenton D. Walker, Stefan Bora
Abstract
Models of parallel processing systems typically assume that one has l servers and jobs are split into an equal number of k = l tasks. This seemingly simple approximation has surprisingly large consequences for the resulting stability and performance bounds. In reality, best practices for modern mapreduce systems indicate that a job's partitioning factor should be much larger than the number of servers available, with some researchers going to far as to advocate for a "tiny tasks" regime, where jobs are split into over 10,000 tasks. In this paper we use recent advances in stochastic network calculus to fundamentally understand the effects of task granularity on parallel systems' scaling, stability, and performance. For the split-merge model, we show that when one allows for tiny tasks, the stability region is actually much better than had previously been concluded. For the single-queue fork-join model, we show that sojourn times quickly approach the optimal case when l "big tasks" are subdivided into k≫ l "tiny tasks". Our results are validated using extensive simulations, and the applicability of the models used is validated by experPiments on an Apache Spark cluster.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Related papers
- Performance and Scaling of Parallel Systems with Blocking Start and/or Departure BarriersBrenton D. Walker, Stefan Bora, Markus FidlerINFOCOM 2022 · 3 citations
- Efficient Microsecond-scale Blind Scheduling with Tiny QuantaZhihong Luo, Sam Son, Dev Bali, Emmanuel Amaro et al.ASPLOS 2024 · 8 citations
- Itoyori: Reconciling Global Address Space and Global Fork-Join Task ParallelismShumpei Shiina, Kenjiro TauraSC 2023 · 6 citations
- QaaD (Query-as-a-Data): Scalable Execution of Massive Number of Small Queries in SparkYeonsu Park, Byungchul Tak, Wook-Shin HanSIGMOD 2023 · 3 citations
- AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning ServingZhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu et al.OSDI 2023 · 211 citations
