User-Defined Operators: Efficiently Integrating Custom Algorithms into Modern Databases
Moritz Sichert, Thomas Neumann
Abstract
In recent years, complex data mining and machine learning algorithms have become more common in data analytics. Several specialized systems exist to evaluate these algorithms on ever-growing data sets, which are built to efficiently execute different types of complex analytics queries. However, using these various systems comes at a price. Moving data out of traditional database systems is often slow as it requires exporting and importing data, which is typically performed using the relatively inefficient CSV format. Additionally, database systems usually offer strong ACID guarantees, which are lost when adding new, external systems. This disadvantage can be detrimental to the consistency of the results. Most data scientists still prefer not to use classical database systems for data analytics. The main reason why RDBMS are not used is that SQL is difficult to work with due to its declarative and set-oriented nature, and is not easily extensible. We present User-Defined Operators (UDOs) as a concept to include custom algorithms into modern query engines. Users can write idiomatic code in the programming language of their choice, which is then directly integrated into existing database systems. We show that our implementation can compete with specialized tools and existing query engines while retaining all beneficial properties of the database system.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b7c1049-bec7-46b0-afe6-e79c44096f7eCited by top-tier papers6
- Declarative Sub-Operators for Universal Data ProcessingMichael Jungmair, Jana GicevaVLDB 2023 · 17 citations
- Hardware-Efficient Data Imputation through DBMS ExtensibilityHubert Mohr-Daurat, Georgios Theodorakis, Holger PirkVLDB 2024 · 2 citations
- Decisionhouse: Prescriptive Analytics in the Data StackMatteo Brucato, Fjodor Kholodkov, Soren Little, Jakob Mayer et al.VLDB 2026 · 2 citations
- FlowLog: Efficient and Extensible Datalog via IncrementalityHangdong Zhao, Zhenghong Yu, Srinag Rao, Simon Frisk et al.VLDB 2026
- The UDFBench Benchmark for General-purpose UDF QueriesYannis Foufoulas, Theoni Palaiologou, Alkis SimitsisVLDB 2025
Builds on6
- Distributed Deep Learning on Data Systems: A Comparative Analysis of ApproachesYuhao Zhang, Frank Mcquillan, Nandish Jayaram, Nikhil Kak et al.VLDB 2021 · 35 citations
- Procedural Extensions of SQL: Understanding their usage in the wildSurabhi Gupta, Karthik RamachandraVLDB 2021 · 34 citations
- One WITH RECURSIVE is Worth Many GOTOsDenis Hirn, Torsten GrustSIGMOD 2021 · 19 citations
- Adaptive Code Generation for Data-Intensive AnalyticsWangda Zhang, Junyoung Kim, Kenneth A. Ross, Eric Sedlar et al.VLDB 2021 · 12 citations
- ParaX: Boosting Deep Learning for Big Data Analytics on Many-Core CPUsLujia Yin, Yiming Zhang, Zhaoning Zhang, Yuxing Peng et al.VLDB 2021 · 9 citations
Related papers
- YeSQL: "You extend SQL" with Rich and Highly Performant User-Defined Functions in Relational DatabasesYannis E. Foufoulas, Alkis Simitsis, Eleftherios Stamatogiannakis, Yannis E. IoannidisVLDB 2022 · 25 citations
- DB4ML - An In-Memory Database Kernel with Machine Learning SupportMatthias Jasny, Tobias Ziegler, Tim Kraska, Uwe Röhm et al.SIGMOD 2020 · 27 citations
- Towards Automatic and Efficient Prediction Query Processing in Analytical DatabaseYuchen Peng, Zhongle Xie, Ke Chen, Gang Chen et al.ICDE 2025 · 3 citations
- PyTond: Efficient Python Data Science on the Shoulders of DatabasesHesam Shahrokhi, Amirali Kaboli, Mahdi Ghorbani, Amir ShaikhhaICDE 2024 · 4 citations
- Database Technology for the Masses: Sub-Operators as First-Class EntitiesMaximilian Bandle, Jana GicevaVLDB 2021 · 19 citations
