Aegis: A Correlation-Based Data Masking Advisor for Data Sharing Ecosystems
Omar Islam Laskar, Fatemeh Ramezani Khozestani, Ishika Nankani, Sohrab Namazi Nia, Senjuti Basu Roy, Kaustubh Beedkar
Abstract
Data-sharing ecosystems connect providers, consumers, and intermediaries to facilitate the exchange and use of data for a wide range of downstream tasks. In sensitive domains such as healthcare, privacy is enforced as a hard constraint--any shared data must satisfy a minimum privacy threshold. However, among all masking configurations that meet this requirement, the utility of the masked data can vary significantly, posing a key challenge: how to efficiently select the optimal configuration that preserves maximum utility. This paper presents A egis , a middleware framework that selects optimal masking configurations for machine learning datasets with features and class labels. A egis incorporates a utility optimizer that minimizes predictive utility deviation --quantifying shifts in feature-label correlations due to masking. Our framework leverages limited data summaries (such as 1D histograms) or none to estimate the feature-label joint distribution, making it suitable for scenarios where raw data is inaccessible due to privacy restrictions. To achieve this, we propose a joint distribution estimator based on iterative proportional fitting, which allows supporting various feature-label correlation quantification methods such as mutual information, chi-square, or g3. Our experimental evaluation of real-world datasets shows that Aegis identifies optimal masking configurations over an order of magnitude faster, while the resulting masked datasets achieve predictive performance on downstream ML tasks on par with baseline approaches and complements privacy anonymization data masking techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 749534fe-6365-40e8-a179-b6fb56b65563Builds on5
- Federated Learning on Non-IID Data Silos: An Experimental StudyQinbin Li, Yiqun Diao, Quan Chen, Bingsheng HeICDE 2022 · 1,110 citations
- Efficient Approximate Algorithms for Empirical Entropy and Mutual InformationXingguang Chen, Sibo WangSIGMOD 2021 · 9 citations
- Fainder: A Fast and Accurate Index for Distribution-Aware Dataset SearchLennart Behme, Sainyam Galhotra, Kaustubh Beedkar, Volker MarklVLDB 2024 · 9 citations
- Disclosure-Compliant Query AnsweringRudi Poepsel Lemaitre, Kaustubh Beedkar, Volker MarklSIGMOD 2025 · 1 citation
- Synthetic Data - Anonymisation Groundhog DayTheresa Stadler, Bristena Oprisanu, Carmela TroncosoUSENIX Security 2022
Related papers
- Not All Features Are Equal: Discovering Essential Features for Preserving Prediction PrivacyFatemehsadat Mireshghallah, Mohammadkazem Taram, Ali Jalali, Ahmed Taha Elthakeb et al.WWW 2021 · 59 citations
- Operationalizing Data Minimization for Privacy-Preserving LLM PromptingJijie Zhou, Niloofar Mireshghallah, Tianshi LiICLR 2026 · 13 citations
- Shielding PII to Prevent Re-identification and Preserve UtilityShuhao Liu, Wenfei Fan, Yijia XuSIGMOD 2026
- Understanding Fairness and Prediction Error through Subspace Decomposition and Influence AnalysisEnze Shi, Pankaj Bhagwat, Zhixian Yang, Linglong Kong et al.NeurIPS 2025
- Learning from Aggregate responses: Instance Level versus Bag Level Loss FunctionsAdel Javanmard, Lin Chen, Vahab Mirrokni, Ashwinkumar Badanidiyuru et al.ICLR 2024 · 3 citations
