Rhyme Native: Efficient Code Generation for Structured and Semi-Structured Workloads
Ran Guo, Tiark Rompf
摘要
Modern data processing spans two worlds: flat relational tables, served by decades of database research producing highly optimized query engines, and nested semi-structured data such as JSON, for which expressive query languages exist but compilation and optimization techniques have been applied far less comprehensively. We ask whether a single query language can express both regimes naturally while compiling to efficient native code. We build on Rhyme, a declarative language whose object-notation syntax mirrors the structure of query results, and contribute on three fronts. We refine Rhyme's semantics for generator binding and missing values, allowing co-iteration, inner/outer joins, and nested-loop traversals to be expressed under different uses of generator symbols. We show that Rhyme's prior dependency-driven loop scheduler can generate incorrect code on hierarchical queries, and present a new scheduler based on finer-grained per-statement constraints that ensures correctness. We introduce a gradual type system and a C code generation backend that emits tag-less, statically typed code and specializes data loading and internal data structures for idiomatic SQL patterns. Our system matches state-of-the-art compiled engines on SQL workloads such as TPC-H and outperforms modern JSON-capable databases and DSLs on JSONBench and other hierarchical queries.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Functional collection programming with semi-ring dictionariesAmir Shaikhha, Mathieu Huot, Jaclyn Smith, Dan OlteanuOOPSLA 2022 · 被引用 31 次
- JSON Tiles: Fast Analytics on Semi-Structured DataDominik Durner, Viktor Leis, Thomas NeumannSIGMOD 2021 · 被引用 28 次
- Rumble: Data Independence for Large Messy Data SetsIngo Müller, Ghislain Fourny, Stefan Irimescu, Can Berker Cikis 等VLDB 2021 · 被引用 12 次
相关 Paper
- Designing an Open Framework for Query Optimization and CompilationMichael Jungmair, André Kohn, Jana GicevaVLDB 2022 · 被引用 45 次
- Dynamic Speculative Optimizations for SQL Compilation in Apache SparkFilippo Schiavio, Daniele Bonetta, Walter BinderVLDB 2020 · 被引用 11 次
- Declarative Sub-Operators for Universal Data ProcessingMichael Jungmair, Jana GicevaVLDB 2023 · 被引用 17 次
- Deductive optimization of relational data storageJohn K. Feser, Sam Madden, Nan Tang, Armando Solar-LezamaOOPSLA 2020 · 被引用 5 次
- Language-Agnostic Integrated Queries in a Managed Polyglot RuntimeFilippo Schiavio, Daniele Bonetta, Walter BinderVLDB 2021 · 被引用 6 次
