Mitigating Structural Overfitting: A Distribution-Aware Rectification Framework for Missing Feature Imputation
Yifan Song, Fenglin Yu, Yihong Luo, Xingjian Tao, Siya Qiu, Kai Han, Jing Tang
Abstract
Incomplete node features are ubiquitous in real-world scenarios such as user profiling and cold-start recommendation, which severely hinders the practical deployment of graph learning systems (e.g., GNNs). Existing solutions typically rely on diffusion-based structural smoothing (e.g., feature propagation) to impute missing values. However, we find that these approaches suffer from structural overfitting, leading to three progressive challenges: 1) performance degradation on disjoint graphs, 2) loss of semantic diversity due to over-smoothing, and 3) feature distribution shift when generalizing to unseen graph structures (inductive tasks). To address these challenges, we introduce the DART framework. It begins by employing Global Structural Augmentation (GSA), which establishes global correlations to bridge disjoint components and extend diffusion coverage. Building upon this, we design a semantic rectifier based on masked autoencoding. This module learns the latent feature manifold to recover natural semantic details. Crucially, we introduce a test-time distribution rectification mechanism that projects structurally biased features back onto the learned manifold during inference, effectively bridging the inductive distribution gap. Furthermore, considering that synthetic masking fails to reflect realworld sparsity, we present a new dataset Sailing collected from voyage records with naturally missing attributes. Extensive experiments on six public datasets and Sailing demonstrate that DART significantly outperforms state-of-the-art methods in both transductive and inductive settings. Our code and dataset are available at https://github.com/yfsong00/DART.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97d05805-add3-481f-9283-b396f098701bCited by top-tier papers1
Ask how each one uses itBuilds on19
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Measuring and Relieving the Over-Smoothing Problem for Graph Neural Networks from the Topological ViewDeli Chen, Yankai Lin, Wei Li, Peng Li et al.AAAI 2020 · 1,353 citations
- GraphSAINT: Graph Sampling Based Inductive Learning MethodHanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava, Rajgopal Kannan et al.ICLR 2020 · 1,155 citations
- GraphMAE: Self-Supervised Masked Graph AutoencodersZhenyu Hou, Xiao Liu, Yukuo Cen, Yuxiao Dong et al.KDD 2022 · 533 citations
- Symmetric Graph Convolutional Autoencoder for Unsupervised Graph Representation LearningJiwoong Park, Minsik Lee, Hyung Jin Chang, Kyuewang Lee et al.ICCV 2019 · 280 citations
Related papers
- Learning the Latent Structure: A Feature-Centric Approach to Graph Data AugmentationYu Song, Zhigang Hua, Yan Xie, Bingheng Li et al.AAAI 2026
- GLAD: Bidirectional Structure-Attribute Alignment via Latent Graph Diffusion ModelsJiankai Zuo, Yu Zhang, Yang Zhang, Zihao Yao et al.ICML 2026
- MIGDiff: Multi-attributes Imputations for Attribute-missing Graphs via Graph Denoising Diffusion ModelYe Liu, Yang Chen, Hongmin CaiAAAI 2026
- Simple yet Effective Diffusion-based Graph Data Augmentation via Complementary Diffusion TransferLonglong Lin, Youan Zhang, Zeli Wang, Xin LuoKDD 2026
- Propagate and Inject: Revisiting Propagation-Based Feature Imputation for Graphs with Partially Observed FeaturesDaeho Um, Sunoh Kim, Jiwoong Park, Jongin Lim et al.ICML 2025
