On the Discrimination Risk of Mean Aggregation Feature Imputation in Graphs
Arjun Subramonian, Kai-Wei Chang, Yizhou Sun
摘要
In human networks, nodes belonging to a marginalized group often have a disproportionate rate of unknown or missing features. This, in conjunction with graph structure and known feature biases, can cause graph feature imputation algorithms to predict values for unknown features that make the marginalized group's feature values more distinct from the the dominant group's feature values than they are in reality. We call this distinction the discrimination risk. We prove that a higher discrimination risk can amplify the unfairness of a machine learning model applied to the imputed data. We then formalize a general graph feature imputation framework called mean aggregation imputation and theoretically and empirically characterize graphs in which applying this framework can yield feature values with a high discrimination risk. We propose a simple algorithm to ensure mean aggregation-imputed features provably have a low discrimination risk, while minimally sacrificing reconstruction error (with respect to the imputation objective). We evaluate the fairness and accuracy of our solution on synthetic and real-world credit networks. Related work Feature imputation Feature imputation algorithms leverage known feature values to predict unknown feature values (and sometimes update known feature values). For example, unknown feature values may be filled as the mean of known values [16] . However, more intricate feature imputation methods have been proposed in the ML, statistics, and epidemiology literature, with popular approaches including matrix completion [32, 33, 34, 35] , nearest neighbors [36, 37] , multiple imputation via conditional models [38, 39, 40] , and causal inference [41, 42] . Notably, while feature imputation may be applied to data with unknown feature values prior to the data being passed to a ML model,
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Aleatoric and Epistemic Discrimination: Fundamental Limits of Fairness InterventionsHao Wang, Luxi He, Rui Gao, Flávio P. CalmonNeurIPS 2023 · 被引用 28 次
- Adapting Fairness Interventions to Missing ValuesRaymond Feng, Flávio P. Calmon, Hao WangNeurIPS 2023 · 被引用 20 次
- Networked Inequality: Preferential Attachment Bias in Graph Neural Network Link PredictionArjun Subramonian, Levent Sagun, Yizhou SunICML 2024 · 被引用 8 次
- Fair Graph Machine Learning under Adversarial Missingness ProcessesDebolina Halder Lina, Arlei SilvaICLR 2026
它引用的顶会 Paper14
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- GraphSAINT: Graph Sampling Based Inductive Learning MethodHanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava, Rajgopal Kannan 等ICLR 2020 · 被引用 1,155 次
- Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image RepresentationsTianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang 等ICCV 2019 · 被引用 469 次
- Fairness without Demographics through Adversarially Reweighted LearningPreethi Lahoti, Alex Beutel, Jilin Chen, Kang Lee 等NeurIPS 2020 · 被引用 406 次
- Handling Missing Data with Graph Representation LearningJiaxuan You, Xiaobai Ma, Daisy Yi Ding, Mykel J. Kochenderfer 等NeurIPS 2020 · 被引用 274 次
相关 Paper
- Fairness without Imputation: A Decision Tree Approach for Fair Prediction with Missing ValuesHaewon Jeong, Hao Wang, Flávio P. CalmonAAAI 2022 · 被引用 48 次
- Divide-Then-Rule: A Cluster-Driven Hierarchical Interpolator for Attribute-Missing GraphsYaowen Hu, Wenxuan Tu, Yue Liu, Miaomiao Li 等ACM MM 2025 · 被引用 2 次
- Assessing Fairness in the Presence of Missing DataYiliang Zhang, Qi LongNeurIPS 2021 · 被引用 51 次
- Confidence-Based Feature Imputation for Graphs with Partially Known FeaturesDaeho Um, Jiwoong Park, Seulki Park, Jin Young ChoiICLR 2023 · 被引用 3 次
- Towards Multiple Missing Values-resistant Unsupervised Graph Anomaly DetectionJiazhen Chen, Xiuqin Liang, Sichao Fu, Zheng Ma 等AAAI 2026
