Efficient Control Flow in Dataflow Systems: When Ease-of-Use Meets High Performance
Gábor E. Gévay, Tilmann Rabl, Sebastian Breß, Lorand Madai-Tahy, Jorge-Arnulfo Quiané-Ruiz, Volker Markl
Abstract
Modern data analysis tasks often involve control flow statements, such as iterations. Common examples are PageRank and K-means. To achieve scalability, developers usually implement data analysis tasks in distributed dataflow systems, such as Spark and Flink. However, for tasks with control flow statements, these systems still either suffer from poor performance or are hard to use. For example, while Flink supports iterations and Spark provides ease-of-use, Flink is hard to use and Spark has poor performance for iterative tasks. As a result, developers typically have to implement different workarounds to run their jobs with control flow statements in an easy and efficient way.We propose Mitos, a system that achieves the best of both worlds: it achieves both high performance and ease-of-use. Mitos uses an intermediate representation that abstracts away specific control flow statements and is able to represent any imperative control flow. This facilitates building the dataflow graph and coordinating the distributed execution of control flow in a way that is not tied to specific control flow constructs. Our experimental evaluation shows that the performance of Mitos is more than one order of magnitude better than systems that launch new dataflow jobs for every iteration step. Remarkably, it is also up to 10.5 times faster than Flink, which has native iteration support, while matching the ease-of-use of Spark.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- Babelfish: Efficient Execution of Polyglot QueriesPhilipp Marian Grulich, Steffen Zeuch, Volker MarklVLDB 2022 · 32 citations
- The Power of Nested Parallelism in Big Data Processing - Hitting Three Flies with One Slap -Gábor E. Gévay, Jorge-Arnulfo Quiané-Ruiz, Volker MarklSIGMOD 2021 · 7 citations
- Redundancy Elimination in Distributed Matrix ComputationZihao Chen, Baokun Han, Chen Xu, Weining Qian et al.SIGMOD 2022 · 4 citations
Related papers
- CrystalPerf: Learning to Characterize the Performance of Dataflow Computation through Code AnalysisHuangshi Tian, Minchen Yu, Wei WangUSENIX ATC 2021 · 2 citations
- Translation of Array-Based Loops to Distributed Data-Parallel ProgramsLeonidas Fegaras, Md Hasanuzzaman NoorVLDB 2020 · 13 citations
- Pasta: A Cost-Based Optimizer for Generating Pipelining Schedules for Dataflow DAGsXiaozhen Liu, Yicong Huang, Xinyuan Lin, Avinash Kumar et al.SIGMOD 2025 · 1 citation
- Kimbap: A Node-Property Map System for Distributed Graph AnalyticsHochan Lee, Roshan Dathathri, Keshav PingaliASPLOS 2024 · 2 citations
- Fries: Fast and Consistent Runtime Reconfiguration in Dataflow Systems with Transactional GuaranteesZuozhi Wang, Shengquan Ni, Avinash Kumar, Chen LiVLDB 2023 · 9 citations
