PIDS: Attribute Decomposition for Improved Compression and Query Performance in Columnar Storage
Hao Jiang, Chunwei Liu, Qi Jin, John Paparrizos, Aaron J. Elmore
摘要
We propose PIDS, Pattern Inference Decomposed Storage, an innovative storage method for decomposing string attributes in columnar stores. Using an unsupervised approach, PIDS identifies common patterns in string attributes from relational databases, and uses the discovered pattern to split each attribute into sub-attributes. First, by storing and encoding each sub-attribute individually, PIDS can achieve a compression ratio comparable to Snappy and Gzip. Second, by decomposing the attribute, PIDS can push down many query operators to sub-attributes, thereby minimizing I/O and potentially expensive comparison operations, resulting in the faster execution of query operators.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Volume Under the Surface: A New Accuracy Evaluation Measure for Time-Series Anomaly DetectionJohn Paparrizos, Paul Boniol, Themis Palpanas, Ruey S. Tsay 等VLDB 2022 · 被引用 171 次
- TSB-UAD: An End-to-End Benchmark Suite for Univariate Time-Series Anomaly DetectionJohn Paparrizos, Yuhao Kang, Paul Boniol, Ruey S. Tsay 等VLDB 2022 · 被引用 138 次
- Decomposed Bounded Floats for Fast Compression and QueriesChunwei Liu, Hao Jiang, John Paparrizos, Aaron J. ElmoreVLDB 2021 · 被引用 65 次
- Good to the Last Bit: Data-Driven Encoding with CodecDBHao Jiang, Chunwei Liu, John Paparrizos, Andrew A. Chien 等SIGMOD 2021 · 被引用 45 次
- Choose Wisely: An Extensive Evaluation of Model Selection for Anomaly Detection in Time SeriesEmmanouil Sylligardos, Paul Boniol, John Paparrizos, Panos E. Trahanias 等VLDB 2023 · 被引用 40 次
相关 Paper
- High-Ratio Compression for Machine-Generated DataJiujing Zhang, Zhitao Shen, Shiyu Yang, Lingkai Meng 等SIGMOD 2024 · 被引用 7 次
- Improving LZ4 for Effective Compression and Efficient QueryZhiheng Liu, Shaoxu SongSIGMOD 2026 · 被引用 1 次
- POLARDB Meets Computational Storage: Efficiently Support Analytical Workloads in Cloud-Native Relational DatabaseWei Cao, Yang Liu, Zhushi Cheng, Ning Zheng 等FAST 2020 · 被引用 140 次
- Accelerating String-Heavy Queries with LLM Token TablesTobias Schmidt, Nicolas Schmitt, Thomas Neumann, Andreas KipfVLDB 2026
- Selection Pushdown in Column Stores using Bit Manipulation InstructionsYinan Li, Jianan Lu, Badrish ChandramouliSIGMOD 2023 · 被引用 15 次
