Data Minimization at Inference Time
Cuong Tran, Ferdinando Fioretto
Abstract
In domains with high stakes such as law, recruitment, and healthcare, learning models frequently rely on sensitive user data for inference, necessitating the complete set of features. This not only poses significant privacy risks for individuals but also demands substantial human effort from organizations to verify information accuracy. This paper asks whether it is necessary to use all input features for accurate predictions at inference time. The paper demonstrates that, in a personalized setting, individuals may only need to disclose a small subset of their features without compromising decision-making accuracy. The paper also provides an efficient sequential algorithm to determine the appropriate attributes for each individual to provide. Evaluations across various learning tasks show that individuals can potentially report as little as 10% of their information while maintaining the same accuracy level as a model that employs the full set of user information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f244d8b-4662-46c7-8b90-3e80c0396acdCited by top-tier papers1
Ask how each one uses itBuilds on4
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Improving model calibration with accuracy versus uncertainty optimizationRanganath Krishnan, Omesh TickooNeurIPS 2020 · 217 citations
- Robust and differentially private mean estimationXiyang Liu, Weihao Kong, Sham M. Kakade, Sewoong OhNeurIPS 2021 · 87 citations
- Auditing Black-Box Prediction Models for Data Minimization ComplianceBashir Rastegarpanah, Krishna P. Gummadi, Mark CrovellaNeurIPS 2021 · 24 citations
Related papers
- DISCO: Dynamic and Invariant Sensitive Channel Obfuscation for Deep Neural NetworksAbhishek Singh, Ayush Chopra, Ethan Garza, Emily Zhang et al.CVPR 2021
- When Machine Learning Gets Personal: Evaluating Prediction and ExplanationLouisa Cornelis, Guillermo Bernardez, Haewon Jeong, Nina MiolaneICLR 2026
- Fair Learning with Private Demographic DataHussein Mozannar, Mesrob I. Ohannessian, Nathan SrebroICML 2020 · 85 citations
- Participatory Personalization in ClassificationHailey Joren, Chirag Nagpal, Katherine A. Heller, Berk UstunNeurIPS 2023 · 7 citations
- Learning Fair Naive Bayes Classifiers by Discovering and Eliminating Discrimination PatternsYooJung Choi, Golnoosh Farnadi, Behrouz Babaki, Guy Van den BroeckAAAI 2020 · 31 citations
