BOS: Bit-Packing with Outlier Separation
Jinzhao Xiao, Zihan Guo, Shaoxu Song
Abstract
Bit-packing serves as the fundamental operator in various data encoding and compression methods. The idea is to use a fixed bit-width to represent all the (processed) values in a sequence. Some extremely large values, known as outliers, obviously amplify the bit-width, and thus lead to wasted bits for most other small values. We notice that not only the large values (upper outliers) but also the small ones (lower outliers) could incur wasted bit-width. In this paper, we propose to store both the upper and lower outliers separately, namely Bit-packing with Outlier Separation (BOS). While the remaining center values have a narrow spread, i.e., condensed bit-width, the separated outliers need some extra cost to denote their positions. The problem is thus how to determine better thresholds for separating the upper and lower outliers, yielding smaller storage cost. Rather than enumerating all the possible values as upper and lower outlier separators, intime, we consider bit-width as the separators, withsearch time. Theoretical analysis illustrates all the possible cases such that the bit-width separation still returns the optimal solution as the value separation, and further leads to an approximate separation strategy with both median and bit-width, intime. Remarkably, our BOS is compatible to any existing compression methods using Bit-packing, and has replaced Bit-packing in Apache IoTDB and Apache TsFile. The extensive experiments on many real-world datasets demonstrate that by replacing Bit-packing with the proposed BOS in various compression methods, the compression ratio is significantly improved from about 2.75 to 3.25.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 18e7a8b8-ef89-4fa8-82bf-29eb96bd3418Builds on6
- Chimp: Efficient Lossless Floating Point Compression for Time Series DatabasesPanagiotis Liakos, Katia Papakonstantinopoulou, Yannis KotidisVLDB 2022 · 76 citations
- Decomposed Bounded Floats for Fast Compression and QueriesChunwei Liu, Hao Jiang, John Paparrizos, Aaron J. ElmoreVLDB 2021 · 65 citations
- Elf: Erasing-based Lossless Floating-Point CompressionRuiyuan Li, Zheng Li, Yi Wu, Chao Chen et al.VLDB 2023 · 44 citations
- Time Series Data Encoding for Efficient Storage: A Comparative Analysis in Apache IoTDBJinzhao Xiao, Yuxiang Huang, Changyu Hu, Shaoxu Song et al.VLDB 2022 · 37 citations
- Frequency Domain Data Encoding in Apache IoTDBHaoyu Wang, Shaoxu SongVLDB 2023 · 17 citations
Related papers
- Sorting Compressed Time SeriesZhiheng Liu, Xingyu Liu, Shaoxu Song, Jianmin WangICDE 2026
- REGER: Reordering Time Series Data for Regression EncodingJinzhao Xiao, Wendi He, Shaoxu Song, Xiangdong Huang et al.ICDE 2024 · 1 citation
- MOST: Model-Based Compression with Outlier Storage for Time Series DataZehai Yang, Shimin ChenSIGMOD 2024 · 9 citations
- Deferred Flushing for Out-of-Order Arrivals in Apache IoTDBXiaojian Zhang, Zhiheng Liu, Shaoxu Song, Xiangdong Huang et al.ICDE 2026
- Improving LZ4 for Effective Compression and Efficient QueryZhiheng Liu, Shaoxu SongSIGMOD 2026 · 1 citation
