PyTond: Efficient Python Data Science on the Shoulders of Databases
Hesam Shahrokhi, Amirali Kaboli, Mahdi Ghorbani, Amir Shaikhha
Abstract
Python data science libraries such as Pandas and NumPy have recently gained immense popularity. Although these libraries are feature-rich and easy to use, their scalability limitations require more robust computational resources. In this paper, we present PyTond, an efficient approach to push the processing of data science workloads down into the database engines that are already known for their big data handling capabilities. Compared to the previous work, by introducing TondIR, our approach can capture a more comprehensive set of workloads and data layouts. Moreover, by doing IR-level optimizations, we generate better SQL code that improves the query processing by the underlying database engine. Our evaluation results show promising performance improvement compared to Python and other alternatives for diverse data science workloads.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c7f3a8b-800c-4049-8605-7bd888da3844Cited by top-tier papers5
- TableRAG: A Retrieval Augmented Generation Framework for Heterogeneous Document ReasoningXiaohan Yu, Pu Jian, Chong ChenEMNLP 2025 · 4 citations
- InferF: Declarative Factorization of AI/ML Inferences over JoinsKanchan Chowdhury, Lixi Zhou, Lulu Xie, Xinwei Fu et al.SIGMOD 2026 · 2 citations
- QURE: AI-Assisted and Automatically Verified UDF InliningTarique Siddiqui, Arnd Christian König, Jiashen Cao, Cong Yan et al.SIGMOD 2025 · 2 citations
- Automated Tensor-Relational Decomposition for Large-Scale Sparse Tensor ComputationYuxin Tang, Zhiyuan Xin, Zhimin Ding, Xinyu Yao et al.VLDB 2026
- PICACHV: Formally Verified Data Use Policy Enforcement for Secure Data AnalyticsHaobin Hiroki Chen, Hongbo Chen, Mingshen Sun, Chenghong Wang et al.USENIX Security 2025
Builds on9
- Towards Scalable Dataframe SystemsDevin Petersohn, William W. Ma, Doris Jung Lin Lee, Stephen Macke et al.VLDB 2020 · 109 citations
- Designing an Open Framework for Query Optimization and CompilationMichael Jungmair, André Kohn, Jana GicevaVLDB 2022 · 45 citations
- Tuplex: Data Science in Python at Native Code SpeedLeonhard F. Spiegelberg, Rahul Yesantharao, Malte Schwarzkopf, Tim KraskaSIGMOD 2021 · 32 citations
- Functional collection programming with semi-ring dictionariesAmir Shaikhha, Mathieu Huot, Jaclyn Smith, Dan OlteanuOOPSLA 2022 · 31 citations
- Optimizing Tensor Programs on Flexible StorageMaximilian Schleich, Amir Shaikhha, Dan SuciuSIGMOD 2023 · 21 citations
Related papers
- HiPy: Extracting High-Level Semantics from Python Code for Data ProcessingMichael Jungmair, Alexis Engelke, Jana GicevaOOPSLA 2024 · 7 citations
- YeSQL: "You extend SQL" with Rich and Highly Performant User-Defined Functions in Relational DatabasesYannis E. Foufoulas, Alkis Simitsis, Eleftherios Stamatogiannakis, Yannis E. IoannidisVLDB 2022 · 25 citations
- Moko: Marrying Python with Big Data SystemsKe Meng, Tao He, Sijie Shen, Lei Wang et al.EuroSys 2025 · 2 citations
- PolyFrame: A Retargetable Query-based Approach to Scaling DataframesPhanwadee Sinthong, Michael J. CareyVLDB 2021 · 8 citations
- Accio: Bolt-on Query FederationXiaoying Wang, Jiannan Wang, Tianzheng Wang, Yong ZhangVLDB 2025
