User-Defined Operators: Efficiently Integrating Custom Algorithms into Modern Databases
Moritz Sichert, Thomas Neumann
摘要
In recent years, complex data mining and machine learning algorithms have become more common in data analytics. Several specialized systems exist to evaluate these algorithms on ever-growing data sets, which are built to efficiently execute different types of complex analytics queries. However, using these various systems comes at a price. Moving data out of traditional database systems is often slow as it requires exporting and importing data, which is typically performed using the relatively inefficient CSV format. Additionally, database systems usually offer strong ACID guarantees, which are lost when adding new, external systems. This disadvantage can be detrimental to the consistency of the results. Most data scientists still prefer not to use classical database systems for data analytics. The main reason why RDBMS are not used is that SQL is difficult to work with due to its declarative and set-oriented nature, and is not easily extensible. We present User-Defined Operators (UDOs) as a concept to include custom algorithms into modern query engines. Users can write idiomatic code in the programming language of their choice, which is then directly integrated into existing database systems. We show that our implementation can compete with specialized tools and existing query engines while retaining all beneficial properties of the database system.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Declarative Sub-Operators for Universal Data ProcessingMichael Jungmair, Jana GicevaVLDB 2023 · 被引用 17 次
- Hardware-Efficient Data Imputation through DBMS ExtensibilityHubert Mohr-Daurat, Georgios Theodorakis, Holger PirkVLDB 2024 · 被引用 2 次
- Decisionhouse: Prescriptive Analytics in the Data StackMatteo Brucato, Fjodor Kholodkov, Soren Little, Jakob Mayer 等VLDB 2026 · 被引用 2 次
- FlowLog: Efficient and Extensible Datalog via IncrementalityHangdong Zhao, Zhenghong Yu, Srinag Rao, Simon Frisk 等VLDB 2026
- The UDFBench Benchmark for General-purpose UDF QueriesYannis Foufoulas, Theoni Palaiologou, Alkis SimitsisVLDB 2025
它引用的顶会 Paper6
- Distributed Deep Learning on Data Systems: A Comparative Analysis of ApproachesYuhao Zhang, Frank Mcquillan, Nandish Jayaram, Nikhil Kak 等VLDB 2021 · 被引用 35 次
- Procedural Extensions of SQL: Understanding their usage in the wildSurabhi Gupta, Karthik RamachandraVLDB 2021 · 被引用 34 次
- One WITH RECURSIVE is Worth Many GOTOsDenis Hirn, Torsten GrustSIGMOD 2021 · 被引用 19 次
- Adaptive Code Generation for Data-Intensive AnalyticsWangda Zhang, Junyoung Kim, Kenneth A. Ross, Eric Sedlar 等VLDB 2021 · 被引用 12 次
- ParaX: Boosting Deep Learning for Big Data Analytics on Many-Core CPUsLujia Yin, Yiming Zhang, Zhaoning Zhang, Yuxing Peng 等VLDB 2021 · 被引用 9 次
相关 Paper
- YeSQL: "You extend SQL" with Rich and Highly Performant User-Defined Functions in Relational DatabasesYannis E. Foufoulas, Alkis Simitsis, Eleftherios Stamatogiannakis, Yannis E. IoannidisVLDB 2022 · 被引用 25 次
- DB4ML - An In-Memory Database Kernel with Machine Learning SupportMatthias Jasny, Tobias Ziegler, Tim Kraska, Uwe Röhm 等SIGMOD 2020 · 被引用 27 次
- Towards Automatic and Efficient Prediction Query Processing in Analytical DatabaseYuchen Peng, Zhongle Xie, Ke Chen, Gang Chen 等ICDE 2025 · 被引用 3 次
- PyTond: Efficient Python Data Science on the Shoulders of DatabasesHesam Shahrokhi, Amirali Kaboli, Mahdi Ghorbani, Amir ShaikhhaICDE 2024 · 被引用 4 次
- Database Technology for the Masses: Sub-Operators as First-Class EntitiesMaximilian Bandle, Jana GicevaVLDB 2021 · 被引用 19 次
