Rhyme Native: Efficient Code Generation for Structured and Semi-Structured Workloads
Ran Guo, Tiark Rompf
Abstract
Modern data processing spans two worlds: flat relational tables, served by decades of database research producing highly optimized query engines, and nested semi-structured data such as JSON, for which expressive query languages exist but compilation and optimization techniques have been applied far less comprehensively. We ask whether a single query language can express both regimes naturally while compiling to efficient native code. We build on Rhyme, a declarative language whose object-notation syntax mirrors the structure of query results, and contribute on three fronts. We refine Rhyme's semantics for generator binding and missing values, allowing co-iteration, inner/outer joins, and nested-loop traversals to be expressed under different uses of generator symbols. We show that Rhyme's prior dependency-driven loop scheduler can generate incorrect code on hierarchical queries, and present a new scheduler based on finer-grained per-statement constraints that ensures correctness. We introduce a gradual type system and a C code generation backend that emits tag-less, statically typed code and specializes data loading and internal data structures for idiomatic SQL patterns. Our system matches state-of-the-art compiled engines on SQL workloads such as TPC-H and outperforms modern JSON-capable databases and DSLs on JSONBench and other hierarchical queries.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on3
- Functional collection programming with semi-ring dictionariesAmir Shaikhha, Mathieu Huot, Jaclyn Smith, Dan OlteanuOOPSLA 2022 · 31 citations
- JSON Tiles: Fast Analytics on Semi-Structured DataDominik Durner, Viktor Leis, Thomas NeumannSIGMOD 2021 · 28 citations
- Rumble: Data Independence for Large Messy Data SetsIngo Müller, Ghislain Fourny, Stefan Irimescu, Can Berker Cikis et al.VLDB 2021 · 12 citations
Related papers
- Designing an Open Framework for Query Optimization and CompilationMichael Jungmair, André Kohn, Jana GicevaVLDB 2022 · 45 citations
- Dynamic Speculative Optimizations for SQL Compilation in Apache SparkFilippo Schiavio, Daniele Bonetta, Walter BinderVLDB 2020 · 11 citations
- Declarative Sub-Operators for Universal Data ProcessingMichael Jungmair, Jana GicevaVLDB 2023 · 17 citations
- Deductive optimization of relational data storageJohn K. Feser, Sam Madden, Nan Tang, Armando Solar-LezamaOOPSLA 2020 · 5 citations
- Language-Agnostic Integrated Queries in a Managed Polyglot RuntimeFilippo Schiavio, Daniele Bonetta, Walter BinderVLDB 2021 · 6 citations
