Stochastic Submodular Data Forgetting
Ramón Rico, Arno Siebes, Yannis Velegrakis
Abstract
Our ability to collect data is rapidly surpassing our ability to store it. As a result, organizations are faced with difficult decisions about which data to retain and which to dispose of. Data forgetting, frames this reduction task as a subset selection exercise. Given a relational dataset 𝐷, a query log 𝑄, and a budget 𝐵, the goal is to find a subset 𝐷 * ⊆ 𝐷 with at most 𝐵 tuples such that it is still possible to compute, based solely on 𝐷 * , approximate answers to the expected query workload. Existing data forgetting routines have substantial limitations. They either offer strong theoretical guarantees but lack scalability due to function evaluation (submodular-based), or achieve scalability by avoiding function evaluation but lack theoretical guarantees (amnesia-based). To bridge the gap between the limitations of submodular and amnesia based methods, we propose IndepDF and DepDF: two data forgetting routines that offer scalability by avoiding function evaluation while maintaining strong theoretical guarantees. Our extensive experimental evaluation on real and synthetic datasets demonstrates that our algorithms are capable of matching the performance of the state-of-the-art submodular-based routines while exhibiting a runtime comparable to that of amnesia-based algorithms. In essence, combining the best traits of both.
• Theory of computation → Data structures and algorithms for data management; • Mathematics of computing → Continuous optimization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on1
Related papers
- Training Data Subset Selection for Regression with Controlled Generalization ErrorDurga Sivasubramanian, Rishabh K. Iyer, Ganesh Ramakrishnan, Abir DeICML 2021 · 25 citations
- Distribution-Level Feature Distancing for Machine Unlearning: Towards a Better Trade-off Between Model Utility and ForgettingDasol Choi, Dongbin NaAAAI 2025 · 9 citations
- Mixed-Privacy Forgetting in Deep NetworksAditya Golatkar, Alessandro Achille, Avinash Ravichandran, Marzia Polito et al.CVPR 2021
- Approximating Opaque Top-k QueriesJiwon Chang, Fatemeh NargesianSIGMOD 2025 · 2 citations
- Computing Views of OWL Ontologies for the Semantic WebJiaqi Li, Xuan Wu, Chang Lu, Wenxing Deng et al.WWW 2021 · 5 citations
