The Price of Tailoring the Index to Your Data: Poisoning Attacks on Learned Index Structures
Evgenios M. Kornaropoulos, Silei Ren, Roberto Tamassia
摘要
The concept of learned index structures relies on the idea that the input-output functionality of a database index can be viewed as a prediction task and, thus, implemented using a machine learning model instead of traditional algorithmic techniques. This novel angle for a decades-old problem has inspired exciting results at the intersection of machine learning and data structures. However, the advantage of learned index structures, i.e., the ability to adjust to the data at hand via the underlying ML-model, can become a disadvantage from a security perspective as it could be exploited.
In this work, we present the first study of data poisoning attacks on learned index structures. Our poisoning approach is different from all previous works since the model under attack is trained on a cumulative distribution function (CDF) and, thus, every injection on the training set has a cascading impact on multiple data values. We formulate the first poisoning attacks on linear regression models trained on a CDF, which is a basic building block of the proposed learned index structures. We generalize our poisoning techniques to attack the advanced two-stage design of learned index structures called recursive model index (RMI), which has been shown to outperform traditional B-Trees. We evaluate our attacks under a variety of parameterizations of the model and show that the error of the RMI increases up to 300× and the error of its second-stage models increases up to 3000×.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Learned Index: A Comprehensive Experimental EvaluationZhaoyan Sun, Xuanhe Zhou, Guoliang LiVLDB 2023 · 被引用 87 次
- PLATON: Top-down R-tree Packing with Learned Partition PolicyJingyi Yang, Gao CongSIGMOD 2024 · 被引用 12 次
- Grafite: Taming Adversarial Queries with Optimal Range FiltersMarco Costa, Paolo Ferragina, Giorgio VinciguerraSIGMOD 2024 · 被引用 11 次
- Algorithmic Complexity Attacks on Dynamic Learned IndexesRui Yang, Evgenios M. Kornaropoulos, Yue ChengVLDB 2024 · 被引用 10 次
- Sieve: A Learned Data-Skipping Index for Data AnalyticsYulai Tong, Jiazhen Liu, Hua Wang, Ke Zhou 等VLDB 2023 · 被引用 10 次
它引用的顶会 Paper14
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning ModelsAhmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang 等NDSS 2019 · 被引用 1,141 次
- Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression LearningMatthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu 等S&P 2018 · 被引用 867 次
- When Does Machine Learning FAIL? Generalized Transferability for Evasion and Poisoning AttacksOctavian Suciu, Radu Marginean, Yigitcan Kaya, Hal Daumé III 等USENIX Security 2018 · 被引用 321 次
- NIC: Detecting Adversarial Samples with Neural Network Invariant CheckingShiqing Ma, Yingqi Liu, Guanhong Tao, Wen-Chuan Lee 等NDSS 2019 · 被引用 283 次
相关 Paper
- Mathematical Foundations of Poisoning Attacks on Linear Regression over Cumulative Distribution FunctionsAtsuki Sato, Martin Aumüller, Yusuke MatsuiSIGMOD 2026
- PACE: Poisoning Attacks on Learned Cardinality EstimationJintao Zhang, Chao Zhang, Guoliang Li, Chengliang ChaiSIGMOD 2024 · 被引用 9 次
- Learned Index with Dynamic Daoyuan Chen, Wuchao Li, Yaliang Li, Bolin Ding 等ICLR 2023
- High Performance or Low Memory? An Updatable Learned Index Framework for Time-Space TradeoffHui Wang, Xin Wang, Jiake Ge, Yunpeng Chai 等SIGMOD 2026
- The Case for Learned In-Memory JoinsIbrahim Sabek, Tim KraskaVLDB 2023 · 被引用 27 次
