Nested Parquet Is Flat, Why Not Use It? How To Scan Nested Data With On-the-Fly Key Generation and Joins
Alice Rey, Maximilian Rieger, Thomas Neumann
摘要
Parquet is the most commonly used file format to store data in a columnar, binary structure. The format also supports storing nested data in this flattened columnar layout. However, many query engines either do not support nested data or process it with substantially worse performance than relational data. In this work, we close this gap and present a new way to leverage relational query engines for nested data that is stored in this flat columnar file format. Specifically, we demonstrate how to process nested Parquet files much more efficiently. Our approach does not store a copy of the data in an internal format but reads directly from the Parquet file. During query computation, the required flat columns are scanned independently and the nesting is reconstructed using joins with on-the-fly generated join keys. Our approach can be easily integrated into existing query engines to support querying nested Parquet files. Furthermore, we achieve orders of magnitude faster analytical query performance than existing solutions, which makes it a valuable addition.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- These Rows Are Made for Sorting and That's Just What We'll DoLaurens Kuiper, Hannes MühleisenICDE 2023 · 被引用 6 次
- NestGPU: Nested Query Processing on GPUSofoklis Floratos, Mengbai Xiao, Hao Wang, Chengxin Guo 等ICDE 2021 · 被引用 17 次
- An Empirical Evaluation of Columnar Storage FormatsXinyu Zeng, Yulong Hui, Jiahong Shen, Andrew Pavlo 等VLDB 2024 · 被引用 59 次
- Selection Pushdown in Column Stores using Bit Manipulation InstructionsYinan Li, Jianan Lu, Badrish ChandramouliSIGMOD 2023 · 被引用 15 次
- A Deep Dive into Common Open Formats for Analytical DBMSsChunwei Liu, Anna Pavlenko, Matteo Interlandi, Brandon HaynesVLDB 2023 · 被引用 23 次
