Lune

VLDB2020Top-tier venue

PIDS: Attribute Decomposition for Improved Compression and Query Performance in Columnar Storage

Hao Jiang, Chunwei Liu, Qi Jin, John Paparrizos, Aaron J. Elmore

2020Year
23Top-tier citations

Abstract

We propose PIDS, Pattern Inference Decomposed Storage, an innovative storage method for decomposing string attributes in columnar stores. Using an unsupervised approach, PIDS identifies common patterns in string attributes from relational databases, and uses the discovered pattern to split each attribute into sub-attributes. First, by storing and encoding each sub-attribute individually, PIDS can achieve a compression ratio comparable to Snappy and Gzip. Second, by decomposing the attribute, PIDS can push down many query operators to sub-attributes, thereby minimizing I/O and potentially expensive comparison operations, resulting in the faster execution of query operators.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 6cadb7b7-ccd9-4e0f-b567-a6d762a4136f

Cited by top-tier papers23

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines