Hardware-Efficient Data Imputation through DBMS Extensibility
Hubert Mohr-Daurat, Georgios Theodorakis, Holger Pirk
摘要
The separation of data and code/queries has served Data Management Systems (DBMSs) well for decades. However, while the resulting soundness and rigidity are the basis for many performance-oriented optimizations, it lacks the flexibility to efficiently support modern data science applications: data cleansing, data ingestion/augmentation or generative models. To support such applications without sacrificing performance, we propose a new logical data model called Homoiconic Collection Processing (HCP). HCP is based on a well-known Meta-Programming concept called Homoiconicity (a unified representation for code and data).
In a DBMS, HCP supports the storage of "classic" relational data but also allows the storage and evaluation of code fragments we refer to as "Homoiconic Expressions". Homoiconic Expressions enable applications such as data imputation directly in the database kernel. Implemented naïvely, such flexibility would come at a prohibitive cost in terms of performance. To make HCP performance-competitive with highly-tuned in-memory DBMSs, we develop a novel storage and processing model called Shape-Wise Microbatching (SWM) and implement it in a system called BOSS. BOSS is performance-competitive with high-performance DBMSs while offering unprecedented extensibility. To demonstrate the extensibility, we implement an extension for impute-and-query workloads: BOSS outperforms state-of-the-art homoiconic runtimes and data imputation systems by two to five orders of magnitude.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Horizon: Scalable Dependency-driven Data CleaningEl Kindi Rezig, Mourad Ouzzani, Walid G. Aref, Ahmed K. Elmagarmid 等VLDB 2021 · 被引用 95 次
- Mind the Gap: An Experimental Evaluation of Imputation of Missing Values Techniques in Time SeriesMourad Khayati, Alberto Lerner, Zakhar Tymchenko, Philippe Cudré-MaurouxVLDB 2020 · 被引用 57 次
- User-Defined Operators: Efficiently Integrating Custom Algorithms into Modern DatabasesMoritz Sichert, Thomas NeumannVLDB 2022 · 被引用 23 次
- ZIP: Lazy Imputation during Query ProcessingYiming Lin, Sharad MehrotraVLDB 2024 · 被引用 4 次
相关 Paper
- BOSS - An Architecture for Database Kernel CompositionHubert Mohr-Daurat, Xuan Sun, Holger PirkVLDB 2024 · 被引用 12 次
- Rapid Data Ingestion through DB-OS Co-designKyungmin Lim, Minseok Yoon, Kihwan Kim, Alan David Fekete 等SIGMOD 2025 · 被引用 1 次
- Incremental Fusion: Unifying Compiled and Vectorized Query ExecutionBenjamin Wagner, André Kohn, Peter Boncz, Viktor LeisICDE 2024 · 被引用 3 次
- Declarative Sub-Operators for Universal Data ProcessingMichael Jungmair, Jana GicevaVLDB 2023 · 被引用 17 次
- Chipmink: Efficient Delta Identification for Massive Object GraphsSupawit Chockchowwat, Sumay Thakurdesai, Zhaoheng Li, Matthew Krafczyk 等VLDB 2026 · 被引用 1 次
