DeepMapping: Learned Data Mapping for Lossless Compression and Efficient Lookup
Lixi Zhou, K. Selçuk Candan, Jia Zou
摘要
Storing tabular data to balance storage and query efficiency is a long-standing research question in the database community. In this work, we argue and show that a novel DeepMapping abstraction, which relies on the impressive memorization capabilities of deep neural networks, can provide better storage cost, better latency, and better run-time memory footprint, all at the same time. Such unique properties may benefit a broad class of use cases in capacity-limited devices. Our proposed DeepMapping abstraction transforms a dataset into multiple key-value mappings and constructs a multi-tasking neural network model that outputs the corresponding values for a given input key. To deal with memorization errors, DeepMapping couples the learned neural network with a lightweight auxiliary data structure capable of correcting mistakes. The auxiliary structure design further enables DeepMapping to efficiently deal with insertions, deletions, and updates even without retraining the mapping. We propose a multi-task search strategy for selecting the hybrid DeepMapping structures (including model architecture and auxiliary structure) with a desirable trade-off among memorization capacity, size, and efficiency. Extensive experiments with a real-world dataset, synthetic and benchmark datasets, including TPC-H and TPC-DS, demonstrated that the DeepMapping approach can better balance the retrieving speed and compression ratio against several cutting-edge competitors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- InferF: Declarative Factorization of AI/ML Inferences over JoinsKanchan Chowdhury, Lixi Zhou, Lulu Xie, Xinwei Fu 等SIGMOD 2026 · 被引用 2 次
- Learned Static Function Data StructuresStefan Hermann, Hans-Peter Lehmann, Giorgio Vinciguerra, Stefan WalzerVLDB 2026
- Time-Varying Vector Field Compression with Preserved Critical Point TrajectoriesMingze Xia, Yuxiao Li, Pu Jiao, Bei Wang 等ICDE 2026
它引用的顶会 Paper14
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- ALEX: An Updatable Adaptive Learned IndexJialin Ding, Umar Farooq Minhas, Jia Yu, Chi Wang 等SIGMOD 2020 · 被引用 274 次
- Learning Multi-Dimensional IndexesVikram Nathan, Jialin Ding, Mohammad Alizadeh, Tim KraskaSIGMOD 2020 · 被引用 180 次
- Tsunami: A Learned Multi-dimensional Index for Correlated Data and Skewed WorkloadsJialin Ding, Vikram Nathan, Mohammad Alizadeh, Tim KraskaVLDB 2021 · 被引用 178 次
- The PGM-index: a fully-dynamic compressed learned index with provable worst-case boundsPaolo Ferragina, Giorgio VinciguerraVLDB 2020 · 被引用 178 次
相关 Paper
- Towards Memory-Efficient Neural Networks via Multi-Level in situ GenerationJiaqi Gu, Hanqing Zhu, Chenghao Feng, Mingjie Liu 等ICCV 2021 · 被引用 4 次
- Effectively Learning Spatial IndicesJianzhong Qi, Guanli Liu, Christian S. Jensen, Lars KulikVLDB 2020 · 被引用 121 次
- DANets: Deep Abstract Networks for Tabular Data Classification and RegressionJintai Chen, Kuanlun Liao, Yao Wan, Danny Z. Chen 等AAAI 2022 · 被引用 82 次
- DeepEverest: Accelerating Declarative Top-K Queries for Deep Neural Network InterpretationDong He, Maureen Daum, Walter Cai, Magdalena BalazinskaVLDB 2022 · 被引用 7 次
- Network Memory Footprint Compression Through Jointly Learnable Codebooks and MappingsEdouard Yvinec, Arnaud Dapogny, Kevin BaillyICLR 2024 · 被引用 2 次
