Tuplex: Data Science in Python at Native Code Speed
Leonhard F. Spiegelberg, Rahul Yesantharao, Malte Schwarzkopf, Tim Kraska
Abstract
Data science pipelines today are either written in Python or contain user-defined functions (UDFs) written in Python. But executing interpreted Python code is slow, and arbitrary Python UDFs cannot be compiled to optimized machine code.
We present Tuplex, the first data analytics framework that just-in-time compiles developers' natural Python UDFs into efficient, end-to-end optimized native code. Tuplex introduces a novel dual-mode execution model that compiles an optimized fast path for the common case, and falls back on slower exception code paths for data that fail to match the fast path's assumptions. Dual-mode execution is crucial to making endto-end optimizing compilation tractable: by focusing on the common case, Tuplex keeps the code simple enough to apply aggressive optimizations. Thanks to dual-mode execution, Tuplex pipelines always complete even if exceptions occur, and Tuplex's post-facto exception handling simplifies debugging.
We evaluate Tuplex with data science pipelines over realworld datasets. Compared to Spark and Dask, Tuplex improves end-to-end pipeline runtime by 5×-109×, and comes within 22% of a hand-optimized C++ baseline. Optimizations enabled by dual-mode processing improve runtime by up to 3×; Tuplex outperforms other Python compilers by 5×; and Tuplex performs well in a distributed setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- YeSQL: "You extend SQL" with Rich and Highly Performant User-Defined Functions in Relational DatabasesYannis E. Foufoulas, Alkis Simitsis, Eleftherios Stamatogiannakis, Yannis E. IoannidisVLDB 2022 · 25 citations
- Predicate Pushdown for Data Science PipelinesCong Yan, Yin Lin, Yeye HeSIGMOD 2023 · 15 citations
- Containerized Execution of UDFs: An Experimental EvaluationKarla Saur, Tara Mirmira, Konstantinos Karanasos, Jesús Camacho-RodríguezVLDB 2022 · 14 citations
- BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler ApproachZhen Zheng, Zaifeng Pan, Dalin Wang, Kai Zhu et al.SIGMOD 2024 · 14 citations
- The Key to Effective UDF Optimization: Before Inlining, First Perform OutliningSamuel Arch, Yuchen Liu, Todd C. Mowry, Jignesh M. Patel et al.VLDB 2025 · 8 citations
Related papers
- SGX-BigMatrix: A Practical Encrypted Data Analytic Framework With Trusted ProcessorsFahad Shaon, Murat Kantarcioglu, Zhiqiang Lin, Latifur KhanCCS 2017 · 83 citations
- Moko: Marrying Python with Big Data SystemsKe Meng, Tao He, Sijie Shen, Lei Wang et al.EuroSys 2025 · 2 citations
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
- QURE: AI-Assisted and Automatically Verified UDF InliningTarique Siddiqui, Arnd Christian König, Jiashen Cao, Cong Yan et al.SIGMOD 2025 · 2 citations
- HiPy: Extracting High-Level Semantics from Python Code for Data ProcessingMichael Jungmair, Alexis Engelke, Jana GicevaOOPSLA 2024 · 7 citations
