JSON Tiles: Fast Analytics on Semi-Structured Data
Dominik Durner, Viktor Leis, Thomas Neumann
Abstract
Developers often prefer flexibility over upfront schema design, making semi-structured data formats such as JSON increasingly popular. Large amounts of JSON data are therefore stored and analyzed by relational database systems. In existing systems, however, JSON's lack of a fixed schema results in slow analytics. In this paper, we present JSON tiles, which, without losing the flexibility of JSON, enables relational systems to perform analytics on JSON data at native speed. JSON tiles automatically detects the most important keys and extracts them transparently - often achieving scan performance similar to columnar storage. At the same time, JSON tiles is capable of handling heterogeneous and changing data. Furthermore, we automatically collect statistics that enable the query optimizer to find good execution plans. Our experimental evaluation compares against state-of-the-art systems and research proposals and shows that our approach is both robust and efficient.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- JEDI: These aren't the JSON documents you're looking for?Thomas Hütter, Nikolaus Augsten, Christoph M. Kirsch, Michael J. Carey et al.SIGMOD 2022 · 14 citations
- BOSS - An Architecture for Database Kernel CompositionHubert Mohr-Daurat, Xuan Sun, Holger PirkVLDB 2024 · 12 citations
- Columnar Formats for Schemaless LSM-based Document StoresWail Y. Alkowaileet, Michael J. CareyVLDB 2022 · 9 citations
- μSlope: High Compression and Fast Search on Semi-Structured LogsRui Wang, Devin Gibson, Kirk Rodrigues, Yu Luo et al.OSDI 2024 · 8 citations
- High-Ratio Compression for Machine-Generated DataJiujing Zhang, Zhitao Shen, Shiyu Yang, Lingkai Meng et al.SIGMOD 2024 · 7 citations
Builds on2
Related papers
- dsJSON: A Distributed SQL JSON ProcessorMajid Saeedan, Ahmed Eldawy, Zhijia ZhaoSIGMOD 2023 · 1 citation
- Nested Parquet Is Flat, Why Not Use It? How To Scan Nested Data With On-the-Fly Key Generation and JoinsAlice Rey, Maximilian Rieger, Thomas NeumannSIGMOD 2025 · 1 citation
- Reducing Ambiguity in Json Schema DiscoveryWilliam Spoth, Oliver Kennedy, Ying Lu, Beda Christoph Hammerschmidt et al.SIGMOD 2021 · 18 citations
- Rumble: Data Independence for Large Messy Data SetsIngo Müller, Ghislain Fourny, Stefan Irimescu, Can Berker Cikis et al.VLDB 2021 · 12 citations
- Dynamic Speculative Optimizations for SQL Compilation in Apache SparkFilippo Schiavio, Daniele Bonetta, Walter BinderVLDB 2020 · 11 citations
