MaSS: Multi-attribute Selective Suppression for Utility-preserving Data Transformation from an Information-theoretic Perspective
Yizhuo Chen, Chun-Fu Chen, Hsiang Hsu, Shaohan Hu, Marco Pistoia, Tarek F. Abdelzaher
Abstract
The growing richness of large-scale datasets has been crucial in driving the rapid advancement and wide adoption of machine learning technologies. The massive collection and usage of data, however, pose an increasing risk for people's private and sensitive information due to either inadvertent mishandling or malicious exploitation. Besides legislative solutions, many technical approaches have been proposed towards data privacy protection. However, they bear various limitations such as leading to degraded data availability and utility, or relying on heuristics and lacking solid theoretical bases. To overcome these limitations, we propose a formal information-theoretic definition for this utility-preserving privacy protection problem, and design a data-driven learnable data transformation framework that is capable of selectively suppressing sensitive attributes from target datasets while preserving the other useful attributes, regardless of whether or not they are known in advance or explicitly annotated for preservation. We provide rigorous theoretical analyses on the operational bounds for our framework, and carry out comprehensive experimental evaluations using datasets of a variety of modalities, including facial images, voice audio clips, and human activity motion sensor signals. Results demonstrate the effectiveness and generalizability of our method under various configurations on a multitude of tasks. Our code is available at https://github.com/ jpmorganchase/MaSS .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- PASS: Private Attributes Protection with Stochastic Data SubstitutionYizhuo Chen, Chun-Fu Chen, Hsiang Hsu, Shaohan Hu et al.ICML 2025
- Subject-level Inference for Realistic Text Anonymization EvaluationMyeong Seok Oh, Dong-Yun Kim, Hanseok Oh, Chaean Kang et al.ACL 2026
- CLOAK: Contrastive Guidance for Latent Diffusion-Based Data ObfuscationXin Yang, Omid ArdakanianUbiComp 2026
Builds on3
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- CIAGAN: Conditional Identity Anonymization Generative Adversarial NetworksMaxim Maximov, Ismail Elezi, Laura Leal-TaixéCVPR 2020
Related papers
- DISCO: Dynamic and Invariant Sensitive Channel Obfuscation for Deep Neural NetworksAbhishek Singh, Ayush Chopra, Ethan Garza, Emily Zhang et al.CVPR 2021
- Adversarial Learning of Privacy-Preserving and Task-Oriented RepresentationsTaihong Xiao, Yi-Hsuan Tsai, Kihyuk Sohn, Manmohan Chandraker et al.AAAI 2020 · 87 citations
- The Fundamental Limits of Least-Privilege LearningTheresa Stadler, Bogdan Kulynych, Michael Gastpar, Nicolas Papernot et al.ICML 2024 · 3 citations
- Trade-offs and Guarantees of Adversarial Representation Learning for Information ObfuscationHan Zhao, Jianfeng Chi, Yuan Tian, Geoffrey J. GordonNeurIPS 2020 · 29 citations
- Task-aware Privacy Preservation for Multi-dimensional DataJiangnan Cheng, Ao Tang, Sandeep ChinchaliICML 2022 · 8 citations
