The Price of Tailoring the Index to Your Data: Poisoning Attacks on Learned Index Structures
Evgenios M. Kornaropoulos, Silei Ren, Roberto Tamassia
Abstract
The concept of learned index structures relies on the idea that the input-output functionality of a database index can be viewed as a prediction task and, thus, implemented using a machine learning model instead of traditional algorithmic techniques. This novel angle for a decades-old problem has inspired exciting results at the intersection of machine learning and data structures. However, the advantage of learned index structures, i.e., the ability to adjust to the data at hand via the underlying ML-model, can become a disadvantage from a security perspective as it could be exploited.
In this work, we present the first study of data poisoning attacks on learned index structures. Our poisoning approach is different from all previous works since the model under attack is trained on a cumulative distribution function (CDF) and, thus, every injection on the training set has a cascading impact on multiple data values. We formulate the first poisoning attacks on linear regression models trained on a CDF, which is a basic building block of the proposed learned index structures. We generalize our poisoning techniques to attack the advanced two-stage design of learned index structures called recursive model index (RMI), which has been shown to outperform traditional B-Trees. We evaluate our attacks under a variety of parameterizations of the model and show that the error of the RMI increases up to 300× and the error of its second-stage models increases up to 3000×.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0ae0802-4e73-4685-9700-e332ff06a359Cited by top-tier papers12
- Learned Index: A Comprehensive Experimental EvaluationZhaoyan Sun, Xuanhe Zhou, Guoliang LiVLDB 2023 · 87 citations
- PLATON: Top-down R-tree Packing with Learned Partition PolicyJingyi Yang, Gao CongSIGMOD 2024 · 12 citations
- Grafite: Taming Adversarial Queries with Optimal Range FiltersMarco Costa, Paolo Ferragina, Giorgio VinciguerraSIGMOD 2024 · 11 citations
- Algorithmic Complexity Attacks on Dynamic Learned IndexesRui Yang, Evgenios M. Kornaropoulos, Yue ChengVLDB 2024 · 10 citations
- Sieve: A Learned Data-Skipping Index for Data AnalyticsYulai Tong, Jiazhen Liu, Hua Wang, Ke Zhou et al.VLDB 2023 · 10 citations
Builds on14
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning ModelsAhmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang et al.NDSS 2019 · 1,141 citations
- Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression LearningMatthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu et al.S&P 2018 · 867 citations
- When Does Machine Learning FAIL? Generalized Transferability for Evasion and Poisoning AttacksOctavian Suciu, Radu Marginean, Yigitcan Kaya, Hal Daumé III et al.USENIX Security 2018 · 321 citations
- NIC: Detecting Adversarial Samples with Neural Network Invariant CheckingShiqing Ma, Yingqi Liu, Guanhong Tao, Wen-Chuan Lee et al.NDSS 2019 · 283 citations
Related papers
- Mathematical Foundations of Poisoning Attacks on Linear Regression over Cumulative Distribution FunctionsAtsuki Sato, Martin Aumüller, Yusuke MatsuiSIGMOD 2026
- PACE: Poisoning Attacks on Learned Cardinality EstimationJintao Zhang, Chao Zhang, Guoliang Li, Chengliang ChaiSIGMOD 2024 · 9 citations
- Learned Index with Dynamic Daoyuan Chen, Wuchao Li, Yaliang Li, Bolin Ding et al.ICLR 2023
- High Performance or Low Memory? An Updatable Learned Index Framework for Time-Space TradeoffHui Wang, Xin Wang, Jiake Ge, Yunpeng Chai et al.SIGMOD 2026
- The Case for Learned In-Memory JoinsIbrahim Sabek, Tim KraskaVLDB 2023 · 27 citations
