CrystalPerf: Learning to Characterize the Performance of Dataflow Computation through Code Analysis
Huangshi Tian, Minchen Yu, Wei Wang
Abstract
Dataflow computation dominates the landscape of big data processing, in which a program is structured as a directed acyclic graph (DAG) of operations. As dataflow computation consumes extensive resources in clusters, making sense of its performance becomes critically important. This, however, can be difficult in practice due to the complexity of DAG execution. In this paper, we propose a new approach that learns to characterize the performance of dataflow computation based on code analysis. Unlike existing performance reasoning techniques, our approach requires no code instrumentation and applies to a wide variety of dataflow frameworks. Our key insight is that the source code of an operation contains learnable syntactic and semantic patterns that reveal how it uses resources. Our approach establishes a performance-resource model that, given a dataflow program, infers automatically how much time each operation has spent on each resource (e.g., CPU, network, disk) from past execution traces and the program source code, using machine learning techniques. We then use the model to predict the program runtime under varying resource configurations. We have implemented our solution as a CLI tool called CrystalPerf. Extensive evaluations in Spark, Flink, and TensorFlow show that CrystalPerf can predict job performance under configuration changes in multiple resources with high accuracy. Real-world case studies further demonstrate that CrystalPerf can accurately detect runtime bottlenecks of DAG jobs, simplifying performance debugging. © 2021 USENIX Annual Technical Conference. All rights reserved.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- ProGraML: A Graph-based Program Representation for Data Flow Analysis and Compiler OptimizationsChris Cummins, Zacharias V. Fisches, Tal Ben-Nun, Torsten Hoefler et al.ICML 2021 · 140 citations
- Learning from the Past: Adaptive Parallelism Tuning for Stream Processing SystemsYuxing Han, Lixiang Chen, Haoyu Wang, Zhanghao Chen et al.ICDE 2025 · 2 citations
- Efficient Control Flow in Dataflow Systems: When Ease-of-Use Meets High PerformanceGábor E. Gévay, Tilmann Rabl, Sebastian Breß, Lorand Madai-Tahy et al.ICDE 2021 · 10 citations
- Pasta: A Cost-Based Optimizer for Generating Pipelining Schedules for Dataflow DAGsXiaozhen Liu, Yicong Huang, Xinyuan Lin, Avinash Kumar et al.SIGMOD 2025 · 1 citation
- IcedTea: Efficient and Responsive Time-Travel Debugging in Dataflow SystemsShengquan Ni, Yicong Huang, Zuozhi Wang, Chen LiVLDB 2025 · 2 citations
