Lune

SIGMOD2021Top-tier venue

Tuplex: Data Science in Python at Native Code Speed

Leonhard F. Spiegelberg, Rahul Yesantharao, Malte Schwarzkopf, Tim Kraska

2021Year
32Citations
12Top-tier citations

Abstract

Data science pipelines today are either written in Python or contain user-defined functions (UDFs) written in Python. But executing interpreted Python code is slow, and arbitrary Python UDFs cannot be compiled to optimized machine code.

We present Tuplex, the first data analytics framework that just-in-time compiles developers' natural Python UDFs into efficient, end-to-end optimized native code. Tuplex introduces a novel dual-mode execution model that compiles an optimized fast path for the common case, and falls back on slower exception code paths for data that fail to match the fast path's assumptions. Dual-mode execution is crucial to making endto-end optimizing compilation tractable: by focusing on the common case, Tuplex keeps the code simple enough to apply aggressive optimizations. Thanks to dual-mode execution, Tuplex pipelines always complete even if exceptions occur, and Tuplex's post-facto exception handling simplifies debugging.

We evaluate Tuplex with data science pipelines over realworld datasets. Compared to Spark and Dask, Tuplex improves end-to-end pipeline runtime by 5×-109×, and comes within 22% of a hand-optimized C++ baseline. Optimizations enabled by dual-mode processing improve runtime by up to 3×; Tuplex outperforms other Python compilers by 5×; and Tuplex performs well in a distributed setting.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers12

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines