On the Discrimination Risk of Mean Aggregation Feature Imputation in Graphs
Arjun Subramonian, Kai-Wei Chang, Yizhou Sun
Abstract
In human networks, nodes belonging to a marginalized group often have a disproportionate rate of unknown or missing features. This, in conjunction with graph structure and known feature biases, can cause graph feature imputation algorithms to predict values for unknown features that make the marginalized group's feature values more distinct from the the dominant group's feature values than they are in reality. We call this distinction the discrimination risk. We prove that a higher discrimination risk can amplify the unfairness of a machine learning model applied to the imputed data. We then formalize a general graph feature imputation framework called mean aggregation imputation and theoretically and empirically characterize graphs in which applying this framework can yield feature values with a high discrimination risk. We propose a simple algorithm to ensure mean aggregation-imputed features provably have a low discrimination risk, while minimally sacrificing reconstruction error (with respect to the imputation objective). We evaluate the fairness and accuracy of our solution on synthetic and real-world credit networks. Related work Feature imputation Feature imputation algorithms leverage known feature values to predict unknown feature values (and sometimes update known feature values). For example, unknown feature values may be filled as the mean of known values [16] . However, more intricate feature imputation methods have been proposed in the ML, statistics, and epidemiology literature, with popular approaches including matrix completion [32, 33, 34, 35] , nearest neighbors [36, 37] , multiple imputation via conditional models [38, 39, 40] , and causal inference [41, 42] . Notably, while feature imputation may be applied to data with unknown feature values prior to the data being passed to a ML model,
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7b2706f-691b-4828-845e-e9ae2dc2211bCited by top-tier papers4
- Aleatoric and Epistemic Discrimination: Fundamental Limits of Fairness InterventionsHao Wang, Luxi He, Rui Gao, Flávio P. CalmonNeurIPS 2023 · 28 citations
- Adapting Fairness Interventions to Missing ValuesRaymond Feng, Flávio P. Calmon, Hao WangNeurIPS 2023 · 20 citations
- Networked Inequality: Preferential Attachment Bias in Graph Neural Network Link PredictionArjun Subramonian, Levent Sagun, Yizhou SunICML 2024 · 8 citations
- Fair Graph Machine Learning under Adversarial Missingness ProcessesDebolina Halder Lina, Arlei SilvaICLR 2026
Builds on14
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- GraphSAINT: Graph Sampling Based Inductive Learning MethodHanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava, Rajgopal Kannan et al.ICLR 2020 · 1,155 citations
- Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image RepresentationsTianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang et al.ICCV 2019 · 469 citations
- Fairness without Demographics through Adversarially Reweighted LearningPreethi Lahoti, Alex Beutel, Jilin Chen, Kang Lee et al.NeurIPS 2020 · 406 citations
- Handling Missing Data with Graph Representation LearningJiaxuan You, Xiaobai Ma, Daisy Yi Ding, Mykel J. Kochenderfer et al.NeurIPS 2020 · 274 citations
Related papers
- Fairness without Imputation: A Decision Tree Approach for Fair Prediction with Missing ValuesHaewon Jeong, Hao Wang, Flávio P. CalmonAAAI 2022 · 48 citations
- Divide-Then-Rule: A Cluster-Driven Hierarchical Interpolator for Attribute-Missing GraphsYaowen Hu, Wenxuan Tu, Yue Liu, Miaomiao Li et al.ACM MM 2025 · 2 citations
- Assessing Fairness in the Presence of Missing DataYiliang Zhang, Qi LongNeurIPS 2021 · 51 citations
- Confidence-Based Feature Imputation for Graphs with Partially Known FeaturesDaeho Um, Jiwoong Park, Seulki Park, Jin Young ChoiICLR 2023 · 3 citations
- Towards Multiple Missing Values-resistant Unsupervised Graph Anomaly DetectionJiazhen Chen, Xiuqin Liang, Sichao Fu, Zheng Ma et al.AAAI 2026
