Good to the Last Bit: Data-Driven Encoding with CodecDB
Hao Jiang, Chunwei Liu, John Paparrizos, Andrew A. Chien, Jihong Ma, Aaron J. Elmore
摘要
Columnar databases rely on specialized encoding schemes to reduce storage requirements. These encodings also enable efficient in-situ data processing. Nevertheless, many existing columnar databases are encoding-oblivious. When storing the data, these systems rely on a global understanding of the dataset or the data types to derive simple rules for encoding selection. Such rule-based selection leads to unsatisfactory performance. Specifically, when performing queries, the systems always decode data into memory, ignoring the possibility of optimizing access to encoded data. We develop CodecDB, an encoding-aware columnar database, to demonstrate the benefit of tightly-coupling the database design with the data encoding schemes. CodecDB chooses in a principled manner the most efficient encoding for a given data column and relies on encoding-aware query operators to optimize access to encoded data. Storage-wise, CodecDB achieves on average 90% accuracy for selecting the best encoding and improves the compression ratio by up to 40% compared to the state-of-the-art encoding selection solution. Query-wise, CodecDB is on average one order of magnitude faster than the latest open-source and commercial columnar databases on the TPC-H benchmark, and on average 3x faster than a recent research project on the Star-Schema Benchmark (SSB).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Volume Under the Surface: A New Accuracy Evaluation Measure for Time-Series Anomaly DetectionJohn Paparrizos, Paul Boniol, Themis Palpanas, Ruey S. Tsay 等VLDB 2022 · 被引用 171 次
- TSB-UAD: An End-to-End Benchmark Suite for Univariate Time-Series Anomaly DetectionJohn Paparrizos, Yuhao Kang, Paul Boniol, Ruey S. Tsay 等VLDB 2022 · 被引用 138 次
- An Empirical Evaluation of Columnar Storage FormatsXinyu Zeng, Yulong Hui, Jiahong Shen, Andrew Pavlo 等VLDB 2024 · 被引用 59 次
- Choose Wisely: An Extensive Evaluation of Model Selection for Anomaly Detection in Time SeriesEmmanouil Sylligardos, Paul Boniol, John Paparrizos, Panos E. Trahanias 等VLDB 2023 · 被引用 40 次
- Fast Adaptive Similarity Search through Variance-Aware QuantizationJohn Paparrizos, Ikraduya Edian, Chunwei Liu, Aaron J. Elmore 等ICDE 2022 · 被引用 34 次
它引用的顶会 Paper5
- Order-Preserving Key Compression for In-Memory Search TreesHuanchen Zhang, Xiaoxuan Liu, David G. Andersen, Michael Kaminsky 等SIGMOD 2020 · 被引用 31 次
- Efficient Query Processing with Optimistically Compressed Hash Tables & Strings in the USSRTim Gubner, Viktor Leis, Peter BonczICDE 2020 · 被引用 2 次
- PIDS: Attribute Decomposition for Improved Compression and Query Performance in Columnar StorageHao Jiang, Chunwei Liu, Qi Jin, John Paparrizos 等VLDB 2020
- FSST: Fast Random Access String CompressionPeter Boncz, Thomas Neumann, Viktor LeisVLDB 2020
- MorphStore: Analytical Query Engine with a Holistic Compression-Enabled Processing ModelPatrick Damme, Annett Ungethüm, Johannes Pietrzyk, Alexander Krause 等VLDB 2020
相关 Paper
- Selection Pushdown in Column Stores using Bit Manipulation InstructionsYinan Li, Jianan Lu, Badrish ChandramouliSIGMOD 2023 · 被引用 15 次
- Robust and Budget-Constrained Encoding Configurations for In-Memory Database SystemsMartin BoissierVLDB 2022 · 被引用 17 次
- LeCo: Lightweight Compression via Learning Serial CorrelationsYihao Liu, Xinyu Zeng, Huanchen ZhangSIGMOD 2024 · 被引用 17 次
- Exploring SIMD Vectorization in Aggregation Pipelines for Encoded IoT DataRui Kang, Shaoxu Song, Jianmin WangICDE 2025
- Time Series Data Encoding for Efficient Storage: A Comparative Analysis in Apache IoTDBJinzhao Xiao, Yuxiang Huang, Changyu Hu, Shaoxu Song 等VLDB 2022 · 被引用 37 次
