Babelfish: Efficient Execution of Polyglot Queries
Philipp Marian Grulich, Steffen Zeuch, Volker Markl
Abstract
Today's users of data processing systems come from different domains, have different levels of expertise, and prefer different programming languages. As a result, analytical workload requirements shifted from relational to polyglot queries involving user-defined functions (UDFs). Although some data processing systems support polyglot queries, they often embed third-party language runtimes. This embedding induces a high performance overhead, as it causes additional data materialization between execution engines.
In this paper, we present Babelfish, a novel data processing engine designed for polyglot queries. Babelfish introduces an intermediate representation that unifies queries from different implementation languages. This enables new, holistic optimizations across operator and language boundaries, e.g., operator fusion and workload specialization. As a result, Babelfish avoids data transfers and enables efficient utilization of hardware resources. Our evaluation shows that Babelfish outperforms state-of-the-art data processing systems by up to one order of magnitude and reaches the performance of handwritten code. With Babelfish, we bridge the performance gap between relational and multi-language UDFs and lay the foundation for the efficient execution of future polyglot workloads.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b5ded81-6f6b-4e92-bf72-540df6a60ce1Cited by top-tier papers7
- Containerized Execution of UDFs: An Experimental EvaluationKarla Saur, Tara Mirmira, Konstantinos Karanasos, Jesús Camacho-RodríguezVLDB 2022 · 14 citations
- In-Situ Cross-Database Query ProcessingHaralampos Gavriilidis, Kaustubh Beedkar, Jorge-Arnulfo Quiané-Ruiz, Volker MarklICDE 2023 · 12 citations
- Nexus: Correlation Discovery over Collections of Spatio-Temporal Tabular DataYue Gong, Sainyam Galhotra, Raul Castro FernandezSIGMOD 2024 · 10 citations
- The Key to Effective UDF Optimization: Before Inlining, First Perform OutliningSamuel Arch, Yuchen Liu, Todd C. Mowry, Jignesh M. Patel et al.VLDB 2025 · 8 citations
- HiPy: Extracting High-Level Semantics from Python Code for Data ProcessingMichael Jungmair, Alexis Engelke, Jana GicevaOOPSLA 2024 · 7 citations
Builds on8
- Grizzly: Efficient Stream Processing Through Adaptive Query CompilationPhilipp M. Grulich, Sebastian Breß, Steffen Zeuch, Jonas Traub et al.SIGMOD 2020 · 41 citations
- LightSaber: Efficient Window Aggregation on Multi-core ProcessorsGeorgios Theodorakis, Alexandros Koliousis, Peter R. Pietzuch, Holger PirkSIGMOD 2020 · 36 citations
- Permutable Compiled Queries: Dynamically Adapting Compiled Queries without RecompilingPrashanth Menon, Amadou Ngom, Todd C. Mowry, Andrew Pavlo et al.VLDB 2021 · 26 citations
- Aggify: Lifting the Curse of Cursor Loops using Custom AggregatesSurabhi Gupta, Sanket Purandare, Karthik RamachandraSIGMOD 2020 · 22 citations
- Adaptive Code Generation for Data-Intensive AnalyticsWangda Zhang, Junyoung Kim, Kenneth A. Ross, Eric Sedlar et al.VLDB 2021 · 12 citations
Related papers
- Incremental Fusion: Unifying Compiled and Vectorized Query ExecutionBenjamin Wagner, André Kohn, Peter Boncz, Viktor LeisICDE 2024 · 3 citations
- Language-Agnostic Integrated Queries in a Managed Polyglot RuntimeFilippo Schiavio, Daniele Bonetta, Walter BinderVLDB 2021 · 6 citations
- Declarative Sub-Operators for Universal Data ProcessingMichael Jungmair, Jana GicevaVLDB 2023 · 17 citations
- BabelFish: Fusing Address Translations for ContainersDimitrios Skarlatos, Umur Darbaz, Bhargava Gopireddy, Nam Sung Kim et al.ISCA 2020 · 18 citations
- Accio: Bolt-on Query FederationXiaoying Wang, Jiannan Wang, Tianzheng Wang, Yong ZhangVLDB 2025
