JSON Tiles: Fast Analytics on Semi-Structured Data
Dominik Durner, Viktor Leis, Thomas Neumann
摘要
Developers often prefer flexibility over upfront schema design, making semi-structured data formats such as JSON increasingly popular. Large amounts of JSON data are therefore stored and analyzed by relational database systems. In existing systems, however, JSON's lack of a fixed schema results in slow analytics. In this paper, we present JSON tiles, which, without losing the flexibility of JSON, enables relational systems to perform analytics on JSON data at native speed. JSON tiles automatically detects the most important keys and extracts them transparently - often achieving scan performance similar to columnar storage. At the same time, JSON tiles is capable of handling heterogeneous and changing data. Furthermore, we automatically collect statistics that enable the query optimizer to find good execution plans. Our experimental evaluation compares against state-of-the-art systems and research proposals and shows that our approach is both robust and efficient.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- JEDI: These aren't the JSON documents you're looking for?Thomas Hütter, Nikolaus Augsten, Christoph M. Kirsch, Michael J. Carey 等SIGMOD 2022 · 被引用 14 次
- BOSS - An Architecture for Database Kernel CompositionHubert Mohr-Daurat, Xuan Sun, Holger PirkVLDB 2024 · 被引用 12 次
- Columnar Formats for Schemaless LSM-based Document StoresWail Y. Alkowaileet, Michael J. CareyVLDB 2022 · 被引用 9 次
- μSlope: High Compression and Fast Search on Semi-Structured LogsRui Wang, Devin Gibson, Kirk Rodrigues, Yu Luo 等OSDI 2024 · 被引用 8 次
- High-Ratio Compression for Machine-Generated DataJiujing Zhang, Zhitao Shen, Shiyu Yang, Lingkai Meng 等SIGMOD 2024 · 被引用 7 次
它引用的顶会 Paper2
相关 Paper
- dsJSON: A Distributed SQL JSON ProcessorMajid Saeedan, Ahmed Eldawy, Zhijia ZhaoSIGMOD 2023 · 被引用 1 次
- Nested Parquet Is Flat, Why Not Use It? How To Scan Nested Data With On-the-Fly Key Generation and JoinsAlice Rey, Maximilian Rieger, Thomas NeumannSIGMOD 2025 · 被引用 1 次
- Reducing Ambiguity in Json Schema DiscoveryWilliam Spoth, Oliver Kennedy, Ying Lu, Beda Christoph Hammerschmidt 等SIGMOD 2021 · 被引用 18 次
- Rumble: Data Independence for Large Messy Data SetsIngo Müller, Ghislain Fourny, Stefan Irimescu, Can Berker Cikis 等VLDB 2021 · 被引用 12 次
- Dynamic Speculative Optimizations for SQL Compilation in Apache SparkFilippo Schiavio, Daniele Bonetta, Walter BinderVLDB 2020 · 被引用 11 次
