Predicate Pushdown for Data Science Pipelines
Cong Yan, Yin Lin, Yeye He
Abstract
Predicate pushdown is a widely adopted query optimization. Existing systems and prior work mostly use pattern-matching rules to decide when a predicate can be pushed through certain operators like join or groupby. However, challenges arise in optimizing for data science pipelines due to the widely used non-relational operators and user-defined functions (UDF) that existing rules would fail to cover. In this paper, we present MagicPush, which decides predicate pushdown using a search-verification approach.MagicPush searches for candidate predicates on pipeline input, which is often not the same as the predicate to be pushed down, and verifies that the pushdown does not change pipeline output with full correctness guarantees. Our evaluation on TPC-H queries and 200 real-world pipelines sampled from GitHub Notebooks shows that MagicPush substantially outperforms a strong baseline that uses a union of rules from prior work - it is able to discover new pushdown opportunities and better optimize 42 real-world pipelines with up to 99% reduction in running time, while discovering all pushdown opportunities found by the existing baseline on remaining cases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 618e180d-fc20-4e1a-b480-e8a5a613847eCited by top-tier papers8
- Data-Semantics-Aware Recommendation of Diverse Pivot TablesWhanhee Cho, Anna FarihaSIGMOD 2026 · 4 citations
- Optimal Predicate Pushdown SynthesisRobert Zhang, Eric Hayden Campbell, Dixin Tang, Isil DilligPLDI 2026 · 1 citation
- Understanding and Optimizing Database Pushdown on Disaggregated StorageHua Zhang, Xiao Li, Yuebin Bai, Ming LiuASPLOS 2026 · 1 citation
- Homomorphism Calculus for User-Defined AggregationsZiteng Wang, Ruijie Fang, Linus Zheng, Dixin Tang et al.OOPSLA 2025
- FaScalSQL: A Fast and Scalable GPU-Accelerated SQL Query Engine for Out-of-Memory TablesChaemin Lim, Suhyun Lee, Jinwoo Choi, Kwanghyun Park et al.ICDE 2026
Builds on17
- Towards Scalable Dataframe SystemsDevin Petersohn, William W. Ma, Doris Jung Lin Lee, Stephen Macke et al.VLDB 2020 · 109 citations
- Quantifying TPC-H Choke Points and Their OptimizationsMarkus Dreseler, Martin Boissier, Tilmann Rabl, Matthias UflackerVLDB 2020 · 91 citations
- Qd-tree: Learning Data Layouts for Big Data AnalyticsZongheng Yang, Badrish Chandramouli, Chi Wang, Johannes Gehrke et al.SIGMOD 2020 · 87 citations
- FlexPushdownDB: Hybrid Pushdown and Caching in a Cloud DBMSYifei Yang, Matt Youill, Matthew E. Woicik, Yizhou Liu et al.VLDB 2021 · 67 citations
- Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science NotebooksCong Yan, Yeye HeSIGMOD 2020 · 64 citations
Related papers
- PLAQUE: Automated Predicate Learning at Query TimeYiming Lin, Sharad MehrotraSIGMOD 2024 · 2 citations
- Optimizing Machine Learning Inference Queries with Correlative Proxy ModelsZhihui Yang, Zuozhi Wang, Yicong Huang, Yao Lu et al.VLDB 2022 · 33 citations
- UPP: Universal Predicate Pushdown to Smart StorageIpoom Jeong, Jinghan Huang, Chuxuan Hu, Dohyun Park et al.ISCA 2025 · 7 citations
- Pushing Data-Induced Predicates Through Joins in Big-Data ClustersLaurel J. Orr, Srikanth Kandula, Surajit ChaudhuriVLDB 2020 · 35 citations
- BLEND: A Unified Data Discovery SystemMahdi Esmailoghli, Christoph Schnell, Renée J. Miller, Ziawasch AbedjanICDE 2025 · 4 citations
