Containerized Execution of UDFs: An Experimental Evaluation
Karla Saur, Tara Mirmira, Konstantinos Karanasos, Jesús Camacho-Rodríguez
摘要
User-defined functions (UDFs) have long been used as the de facto way to extend the capabilities of data management systems. However, they are restricted to the specificities of each DBMS, and recent demands for advanced analytics have increased the need for complex UDFs that may require execution of arbitrary computation written in any programming language, management of library dependencies, portability across environments and engines, and resource isolation. These requirements go beyond what traditional UDFs were designed for, and have given rise to containerized UDFs that enable encapsulation and portability. However, this approach is nascent and can result in significant performance penalties and usability issues. In this paper, we present the first study that spans all stages of containerized UDFs' life cycle, performance bottlenecks in their execution, and extensibility to support different engines. Our experiments show that the performance of containerized UDF execution can be greatly affected by system design choices and that there are many trade-offs to consider. For example, regarding the method of communication with the containerized UDF, we show that binary-based implementations minimize overheads and are more than 2.4x faster than widely used text-based ones. Adopting a newer general-purpose communication method such as Arrow Flight can improve performance dramatically, causing a minimal 10% slowdown compared to non-containerized UDFs. Additionally, containerized UDF start times vary wildly due to program size and complexity, from .07s to 7s in our experiments. Our insights can help DBMS developers make appropriate choices based on individual use cases when designing their systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Biathlon: Harnessing Model Resilience for Accelerating ML Inference PipelinesChaokun Chang, Eric Lo, Chunxiao YeVLDB 2024 · 被引用 5 次
- AnyBlox: A Framework for Self-Decoding DatasetsMateusz Gienieczko, Maximilian Kuschewski, Thomas Neumann, Viktor Leis 等VLDB 2025 · 被引用 3 次
- Unlocking True Elasticity for the Cloud-Native Era with DandelionTom Kuchler, Pinghe Li, Yazhuo Zhang, Lazar Cvetkovic 等SOSP 2025 · 被引用 1 次
它引用的顶会 Paper6
- Faasm: Lightweight Isolation for Efficient Stateful Serverless ComputingSimon Shillaker, Peter R. PietzuchUSENIX ATC 2020 · 被引用 382 次
- Firecracker: Lightweight Virtualization for Serverless ApplicationsAlexandru Agache, Marc Brooker, Alexandra Iordache, Anthony Liguori 等NSDI 2020 · 被引用 197 次
- End-to-end Optimization of Machine Learning Prediction QueriesKwanghyun Park, Karla Saur, Dalitso Banda, Rathijit Sen 等SIGMOD 2022 · 被引用 50 次
- Babelfish: Efficient Execution of Polyglot QueriesPhilipp Marian Grulich, Steffen Zeuch, Volker MarklVLDB 2022 · 被引用 32 次
- Tuplex: Data Science in Python at Native Code SpeedLeonhard F. Spiegelberg, Rahul Yesantharao, Malte Schwarzkopf, Tim KraskaSIGMOD 2021 · 被引用 32 次
相关 Paper
- WAF: An Efficient WebAssembly-Based Execution Environment for User-Defined FunctionsZhuo Huang, Hao Fan, Junhui Peng, Qi Wu 等ICDE 2025 · 被引用 1 次
- The UDFBench Benchmark for General-purpose UDF QueriesYannis Foufoulas, Theoni Palaiologou, Alkis SimitsisVLDB 2025
- YeSQL: "You extend SQL" with Rich and Highly Performant User-Defined Functions in Relational DatabasesYannis E. Foufoulas, Alkis Simitsis, Eleftherios Stamatogiannakis, Yannis E. IoannidisVLDB 2022 · 被引用 25 次
- The Key to Effective UDF Optimization: Before Inlining, First Perform OutliningSamuel Arch, Yuchen Liu, Todd C. Mowry, Jignesh M. Patel 等VLDB 2025 · 被引用 8 次
- SQL Engines Excel at the Execution of Imperative ProgramsTim Fischer, Denis Hirn, Torsten GrustVLDB 2024 · 被引用 3 次
