HiPy: Extracting High-Level Semantics from Python Code for Data Processing
Michael Jungmair, Alexis Engelke, Jana Giceva
Abstract
Data science workloads frequently include Python code, but Python's dynamic nature makes efficient execution hard. Traditional approaches either treat Python as a black box, missing out on optimization potential, or are limited to a narrow domain. However, a deep and efficient integration of user-defined Python code into data processing systems requires extracting the semantics of the entire Python code.
In this paper, we propose a novel approach for extracting the high-level semantics by transforming general Python functions into program generators that generate a statically-typed IR when executed. The extracted IR then allows for high-level, domain-specific optimizations and the generation of efficient C++ code. With our prototype implementation, HiPy, we achieve single-threaded speedups of 2-20x for many workloads. Furthermore, HiPy is also capable of accelerating Python code in other domains like numerical data, where it can sometimes even outperform specialized compilers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 912cda79-20fb-4387-8d53-d07d4147aa4fCited by top-tier papers3
- Towards Designing Future-Proof Data Processing SystemsMichael Jungmair, Jana GicevaVLDB 2025 · 1 citation
- TPCx-AI under the Microscope: A Benchmarking Debt AnalysisIlin Tolovski, Philipp Hildebrandt, Khuzaima Daudjee, Tilmann RablVLDB 2026
- Homomorphism Calculus for User-Defined AggregationsZiteng Wang, Ruijie Fang, Linus Zheng, Dixin Tang et al.OOPSLA 2025
Builds on6
- Babelfish: Efficient Execution of Polyglot QueriesPhilipp Marian Grulich, Steffen Zeuch, Volker MarklVLDB 2022 · 32 citations
- Tuplex: Data Science in Python at Native Code SpeedLeonhard F. Spiegelberg, Rahul Yesantharao, Malte Schwarzkopf, Tim KraskaSIGMOD 2021 · 32 citations
- SOAR: A Synthesis Approach for Data Science API RefactoringAnsong Ni, Daniel Ramos, Aidan Z. H. Yang, Inês Lynce et al.ICSE 2021 · 27 citations
- YeSQL: "You extend SQL" with Rich and Highly Performant User-Defined Functions in Relational DatabasesYannis E. Foufoulas, Alkis Simitsis, Eleftherios Stamatogiannakis, Yannis E. IoannidisVLDB 2022 · 25 citations
- Flexible Rule-Based Decomposition and Metadata Independence in Modin: A Parallel Dataframe SystemDevin Petersohn, Dixin Tang, Rehan Sohail Durrani, Areg Melik-Adamyan et al.VLDB 2022 · 22 citations
Related papers
- PyTond: Efficient Python Data Science on the Shoulders of DatabasesHesam Shahrokhi, Amirali Kaboli, Mahdi Ghorbani, Amir ShaikhhaICDE 2024 · 4 citations
- HARP: holistic analysis for refactoring Python-based analytics programsWeijie Zhou, Yue Zhao, Guoqiang Zhang, Xipeng ShenICSE 2020 · 10 citations
- Concrete Type Inference for Code Optimization using Machine Learning with SMT SolvingFangke Ye, Jisheng Zhao, Jun Shirako, Vivek SarkarOOPSLA 2023 · 5 citations
- Dynamic Generation of Python Bindings for HPC KernelsSteven Zhu, Nader Al Awar, Mattan Erez, Milos GligoricASE 2021 · 3 citations
- Productivity, portability, performance: data-centric PythonAlexandros Nikolaos Ziogas, Timo Schneider, Tal Ben-Nun, Alexandru Calotoiu et al.SC 2021 · 32 citations
