Predicate Pushdown for Data Science Pipelines
Cong Yan, Yin Lin, Yeye He
摘要
Predicate pushdown is a widely adopted query optimization. Existing systems and prior work mostly use pattern-matching rules to decide when a predicate can be pushed through certain operators like join or groupby. However, challenges arise in optimizing for data science pipelines due to the widely used non-relational operators and user-defined functions (UDF) that existing rules would fail to cover. In this paper, we present MagicPush, which decides predicate pushdown using a search-verification approach.MagicPush searches for candidate predicates on pipeline input, which is often not the same as the predicate to be pushed down, and verifies that the pushdown does not change pipeline output with full correctness guarantees. Our evaluation on TPC-H queries and 200 real-world pipelines sampled from GitHub Notebooks shows that MagicPush substantially outperforms a strong baseline that uses a union of rules from prior work - it is able to discover new pushdown opportunities and better optimize 42 real-world pipelines with up to 99% reduction in running time, while discovering all pushdown opportunities found by the existing baseline on remaining cases.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Data-Semantics-Aware Recommendation of Diverse Pivot TablesWhanhee Cho, Anna FarihaSIGMOD 2026 · 被引用 4 次
- Optimal Predicate Pushdown SynthesisRobert Zhang, Eric Hayden Campbell, Dixin Tang, Isil DilligPLDI 2026 · 被引用 1 次
- Understanding and Optimizing Database Pushdown on Disaggregated StorageHua Zhang, Xiao Li, Yuebin Bai, Ming LiuASPLOS 2026 · 被引用 1 次
- Homomorphism Calculus for User-Defined AggregationsZiteng Wang, Ruijie Fang, Linus Zheng, Dixin Tang 等OOPSLA 2025
- FaScalSQL: A Fast and Scalable GPU-Accelerated SQL Query Engine for Out-of-Memory TablesChaemin Lim, Suhyun Lee, Jinwoo Choi, Kwanghyun Park 等ICDE 2026
它引用的顶会 Paper17
- Towards Scalable Dataframe SystemsDevin Petersohn, William W. Ma, Doris Jung Lin Lee, Stephen Macke 等VLDB 2020 · 被引用 109 次
- Quantifying TPC-H Choke Points and Their OptimizationsMarkus Dreseler, Martin Boissier, Tilmann Rabl, Matthias UflackerVLDB 2020 · 被引用 91 次
- Qd-tree: Learning Data Layouts for Big Data AnalyticsZongheng Yang, Badrish Chandramouli, Chi Wang, Johannes Gehrke 等SIGMOD 2020 · 被引用 87 次
- FlexPushdownDB: Hybrid Pushdown and Caching in a Cloud DBMSYifei Yang, Matt Youill, Matthew E. Woicik, Yizhou Liu 等VLDB 2021 · 被引用 67 次
- Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science NotebooksCong Yan, Yeye HeSIGMOD 2020 · 被引用 64 次
相关 Paper
- PLAQUE: Automated Predicate Learning at Query TimeYiming Lin, Sharad MehrotraSIGMOD 2024 · 被引用 2 次
- Optimizing Machine Learning Inference Queries with Correlative Proxy ModelsZhihui Yang, Zuozhi Wang, Yicong Huang, Yao Lu 等VLDB 2022 · 被引用 33 次
- UPP: Universal Predicate Pushdown to Smart StorageIpoom Jeong, Jinghan Huang, Chuxuan Hu, Dohyun Park 等ISCA 2025 · 被引用 7 次
- Pushing Data-Induced Predicates Through Joins in Big-Data ClustersLaurel J. Orr, Srikanth Kandula, Surajit ChaudhuriVLDB 2020 · 被引用 35 次
- BLEND: A Unified Data Discovery SystemMahdi Esmailoghli, Christoph Schnell, Renée J. Miller, Ziawasch AbedjanICDE 2025 · 被引用 4 次
