PyTond: Efficient Python Data Science on the Shoulders of Databases
Hesam Shahrokhi, Amirali Kaboli, Mahdi Ghorbani, Amir Shaikhha
摘要
Python data science libraries such as Pandas and NumPy have recently gained immense popularity. Although these libraries are feature-rich and easy to use, their scalability limitations require more robust computational resources. In this paper, we present PyTond, an efficient approach to push the processing of data science workloads down into the database engines that are already known for their big data handling capabilities. Compared to the previous work, by introducing TondIR, our approach can capture a more comprehensive set of workloads and data layouts. Moreover, by doing IR-level optimizations, we generate better SQL code that improves the query processing by the underlying database engine. Our evaluation results show promising performance improvement compared to Python and other alternatives for diverse data science workloads.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- TableRAG: A Retrieval Augmented Generation Framework for Heterogeneous Document ReasoningXiaohan Yu, Pu Jian, Chong ChenEMNLP 2025 · 被引用 4 次
- InferF: Declarative Factorization of AI/ML Inferences over JoinsKanchan Chowdhury, Lixi Zhou, Lulu Xie, Xinwei Fu 等SIGMOD 2026 · 被引用 2 次
- QURE: AI-Assisted and Automatically Verified UDF InliningTarique Siddiqui, Arnd Christian König, Jiashen Cao, Cong Yan 等SIGMOD 2025 · 被引用 2 次
- Automated Tensor-Relational Decomposition for Large-Scale Sparse Tensor ComputationYuxin Tang, Zhiyuan Xin, Zhimin Ding, Xinyu Yao 等VLDB 2026
- PICACHV: Formally Verified Data Use Policy Enforcement for Secure Data AnalyticsHaobin Hiroki Chen, Hongbo Chen, Mingshen Sun, Chenghong Wang 等USENIX Security 2025
它引用的顶会 Paper9
- Towards Scalable Dataframe SystemsDevin Petersohn, William W. Ma, Doris Jung Lin Lee, Stephen Macke 等VLDB 2020 · 被引用 109 次
- Designing an Open Framework for Query Optimization and CompilationMichael Jungmair, André Kohn, Jana GicevaVLDB 2022 · 被引用 45 次
- Tuplex: Data Science in Python at Native Code SpeedLeonhard F. Spiegelberg, Rahul Yesantharao, Malte Schwarzkopf, Tim KraskaSIGMOD 2021 · 被引用 32 次
- Functional collection programming with semi-ring dictionariesAmir Shaikhha, Mathieu Huot, Jaclyn Smith, Dan OlteanuOOPSLA 2022 · 被引用 31 次
- Optimizing Tensor Programs on Flexible StorageMaximilian Schleich, Amir Shaikhha, Dan SuciuSIGMOD 2023 · 被引用 21 次
相关 Paper
- HiPy: Extracting High-Level Semantics from Python Code for Data ProcessingMichael Jungmair, Alexis Engelke, Jana GicevaOOPSLA 2024 · 被引用 7 次
- YeSQL: "You extend SQL" with Rich and Highly Performant User-Defined Functions in Relational DatabasesYannis E. Foufoulas, Alkis Simitsis, Eleftherios Stamatogiannakis, Yannis E. IoannidisVLDB 2022 · 被引用 25 次
- Moko: Marrying Python with Big Data SystemsKe Meng, Tao He, Sijie Shen, Lei Wang 等EuroSys 2025 · 被引用 2 次
- PolyFrame: A Retargetable Query-based Approach to Scaling DataframesPhanwadee Sinthong, Michael J. CareyVLDB 2021 · 被引用 8 次
- Accio: Bolt-on Query FederationXiaoying Wang, Jiannan Wang, Tianzheng Wang, Yong ZhangVLDB 2025
