DeepSqueeze: Deep Semantic Compression for Tabular Data
Amir Ilkhechi, Andrew Crotty, Alex Galakatos, Yicong Mao, Grace Fan, Xiran Shi, Ugur Çetintemel
摘要
With the rapid proliferation of large datasets, efficient data compression has become more important than ever. Columnar compression techniques (e.g., dictionary encoding, run-length encoding, delta encoding) have proved highly effective for tabular data, but they typically compress individual columns without considering potential relationships among columns, such as functional dependencies and correlations. Semantic compression techniques, on the other hand, are designed to leverage such relationships to store only a subset of the columns necessary to infer the others, but existing approaches cannot effectively identify complex relationships across more than a few columns at a time. We propose DeepSqueeze, a novel semantic compression framework that can efficiently capture these complex relationships within tabular data by using autoencoders to map tuples to a lower-dimensional representation. DeepSqueeze also supports guaranteed error bounds for lossy compression of numerical data and works in conjunction with common columnar compression formats. Our experimental evaluation uses real-world datasets to demonstrate that DeepSqueeze can achieve over a 4x size reduction compared to state-of-the-art alternatives.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- LeCo: Lightweight Compression via Learning Serial CorrelationsYihao Liu, Xinyu Zeng, Huanchen ZhangSIGMOD 2024 · 被引用 17 次
- QCore: Data-Efficient, On-Device Continual Calibration for Quantized ModelsDavid Campos, Bin Yang, Tung Kieu, Miao Zhang 等VLDB 2024 · 被引用 11 次
- High-Ratio Compression for Machine-Generated DataJiujing Zhang, Zhitao Shen, Shiyu Yang, Lingkai Meng 等SIGMOD 2024 · 被引用 7 次
- DeepMapping: Learned Data Mapping for Lossless Compression and Efficient LookupLixi Zhou, K. Selçuk Candan, Jia ZouICDE 2024 · 被引用 6 次
- Generating Succinct Descriptions of Database Schemata for Cost-Efficient Prompting of Large Language ModelsImmanuel TrummerVLDB 2024 · 被引用 5 次
相关 Paper
- Blitzcrank: Fast Semantic Compression for In-memory Online Transaction ProcessingYiming Qiao, Yihan Gao, Huanchen ZhangVLDB 2024 · 被引用 3 次
- Video Compression With Rate-Distortion AutoencodersAmirHossein Habibian, Ties van Rozendaal, Jakub M. Tomczak, Taco CohenICCV 2019 · 被引用 233 次
- Compressing Tabular Data via Latent Variable EstimationAndrea Montanari, Eric WeinerICML 2023
- Good to the Last Bit: Data-Driven Encoding with CodecDBHao Jiang, Chunwei Liu, John Paparrizos, Andrew A. Chien 等SIGMOD 2021 · 被引用 45 次
- Robust and Budget-Constrained Encoding Configurations for In-Memory Database SystemsMartin BoissierVLDB 2022 · 被引用 17 次
