LeCo: Lightweight Compression via Learning Serial Correlations
Yihao Liu, Xinyu Zeng, Huanchen Zhang
Abstract
Lightweight data compression is a key technique that allows column stores to exhibit superior performance for analytical queries. Despite a comprehensive study on dictionary-based encodings to approach Shannon's entropy, few prior works have systematically exploited the serial correlation in a column for compression. In this paper, we propose LeCo (i.e., Learned Compression), a framework that uses machine learning to remove the serial redundancy in a value sequence automatically to achieve an outstanding compression ratio and decompression performance simultaneously. LeCo presents a general approach to this end, making existing (ad-hoc) algorithms such as Frame-of-Reference (FOR), Delta Encoding, and Run-Length Encoding (RLE) special cases under our framework. Our microbenchmark with three synthetic and eight real-world data sets shows that a prototype of LeCo achieves a Pareto improvement on both compression ratio and random access speed over the existing solutions. When integrating LeCo into widely-used applications, we observe up to 5.2× speed up in a data analytical query in the Arrow columnar execution engine, and a 16% increase in RocksDB 's throughput.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- An Empirical Evaluation of Columnar Storage FormatsXinyu Zeng, Yulong Hui, Jiahong Shen, Andrew Pavlo et al.VLDB 2024 · 59 citations
- Making In-Memory Learned Indexes Efficient on DiskJiaoyi Zhang, Kai Su, Huanchen ZhangSIGMOD 2024 · 18 citations
- F3: The Open-Source Data File Format for the FutureXinyu Zeng, Ruijun Meng, Martin Prammer, Wes McKinney et al.SIGMOD 2026 · 10 citations
- The FastLanes File FormatAzim Afroozeh, Peter BonczVLDB 2025 · 9 citations
- Learned Compression of Nonlinear Time Series with Random AccessAndrea Guerra, Giorgio Vinciguerra, Antonio Boffa, Paolo FerraginaICDE 2025 · 5 citations
Builds on11
- ALEX: An Updatable Adaptive Learned IndexJialin Ding, Umar Farooq Minhas, Jia Yu, Chi Wang et al.SIGMOD 2020 · 274 citations
- The PGM-index: a fully-dynamic compressed learned index with provable worst-case boundsPaolo Ferragina, Giorgio VinciguerraVLDB 2020 · 178 citations
- SAND: Streaming Subsequence Anomaly DetectionPaul Boniol, John Paparrizos, Themis Palpanas, Michael J. FranklinVLDB 2021 · 128 citations
- FINEdex: A Fine-grained Learned Index Scheme for Scalable and Concurrent Memory SystemsPengfei Li, Yu Hua, Jingnan Jia, Pengfei ZuoVLDB 2022 · 97 citations
- Good to the Last Bit: Data-Driven Encoding with CodecDBHao Jiang, Chunwei Liu, John Paparrizos, Andrew A. Chien et al.SIGMOD 2021 · 45 citations
Related papers
- LICO: An SIMD-Aware High-Performance Learned Inverted Index Compression FrameworkXianyu Zhu, Qiyu Liu, Guangyi Zhang, Zhibing Sha et al.SIGMOD 2026
- DeepSqueeze: Deep Semantic Compression for Tabular DataAmir Ilkhechi, Andrew Crotty, Alex Galakatos, Yicong Mao et al.SIGMOD 2020 · 30 citations
- Adaptive Compression for Fast Scans on String ColumnsYannis Foufoulas, Lefteris Sidirourgos, Elefterios Stamatogiannakis, Yannis E. IoannidisSIGMOD 2021 · 6 citations
- Data Chunk Compaction in Vectorized ExecutionYiming Qiao, Huanchen ZhangSIGMOD 2025 · 2 citations
- Improving LZ4 for Effective Compression and Efficient QueryZhiheng Liu, Shaoxu SongSIGMOD 2026 · 1 citation
