Tuplex: Data Science in Python at Native Code Speed
Leonhard F. Spiegelberg, Rahul Yesantharao, Malte Schwarzkopf, Tim Kraska
摘要
Data science pipelines today are either written in Python or contain user-defined functions (UDFs) written in Python. But executing interpreted Python code is slow, and arbitrary Python UDFs cannot be compiled to optimized machine code.
We present Tuplex, the first data analytics framework that just-in-time compiles developers' natural Python UDFs into efficient, end-to-end optimized native code. Tuplex introduces a novel dual-mode execution model that compiles an optimized fast path for the common case, and falls back on slower exception code paths for data that fail to match the fast path's assumptions. Dual-mode execution is crucial to making endto-end optimizing compilation tractable: by focusing on the common case, Tuplex keeps the code simple enough to apply aggressive optimizations. Thanks to dual-mode execution, Tuplex pipelines always complete even if exceptions occur, and Tuplex's post-facto exception handling simplifies debugging.
We evaluate Tuplex with data science pipelines over realworld datasets. Compared to Spark and Dask, Tuplex improves end-to-end pipeline runtime by 5×-109×, and comes within 22% of a hand-optimized C++ baseline. Optimizations enabled by dual-mode processing improve runtime by up to 3×; Tuplex outperforms other Python compilers by 5×; and Tuplex performs well in a distributed setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- YeSQL: "You extend SQL" with Rich and Highly Performant User-Defined Functions in Relational DatabasesYannis E. Foufoulas, Alkis Simitsis, Eleftherios Stamatogiannakis, Yannis E. IoannidisVLDB 2022 · 被引用 25 次
- Predicate Pushdown for Data Science PipelinesCong Yan, Yin Lin, Yeye HeSIGMOD 2023 · 被引用 15 次
- Containerized Execution of UDFs: An Experimental EvaluationKarla Saur, Tara Mirmira, Konstantinos Karanasos, Jesús Camacho-RodríguezVLDB 2022 · 被引用 14 次
- BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler ApproachZhen Zheng, Zaifeng Pan, Dalin Wang, Kai Zhu 等SIGMOD 2024 · 被引用 14 次
- The Key to Effective UDF Optimization: Before Inlining, First Perform OutliningSamuel Arch, Yuchen Liu, Todd C. Mowry, Jignesh M. Patel 等VLDB 2025 · 被引用 8 次
相关 Paper
- SGX-BigMatrix: A Practical Encrypted Data Analytic Framework With Trusted ProcessorsFahad Shaon, Murat Kantarcioglu, Zhiqiang Lin, Latifur KhanCCS 2017 · 被引用 83 次
- Moko: Marrying Python with Big Data SystemsKe Meng, Tao He, Sijie Shen, Lei Wang 等EuroSys 2025 · 被引用 2 次
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein 等ASPLOS 2024 · 被引用 693 次
- QURE: AI-Assisted and Automatically Verified UDF InliningTarique Siddiqui, Arnd Christian König, Jiashen Cao, Cong Yan 等SIGMOD 2025 · 被引用 2 次
- HiPy: Extracting High-Level Semantics from Python Code for Data ProcessingMichael Jungmair, Alexis Engelke, Jana GicevaOOPSLA 2024 · 被引用 7 次
