GraphCleaner: Detecting Mislabelled Samples in Popular Graph Learning Benchmarks
Yuwen Li, Miao Xiong, Bryan Hooi
Abstract
Label errors have been found to be prevalent in popular text, vision, and audio datasets, which heavily influence the safe development and evaluation of machine learning algorithms. Despite increasing efforts towards improving the quality of generic data types, such as images and texts, the problem of mislabel detection in graph data remains underexplored. To bridge the gap, we explore mislabelling issues in popular real-world graph datasets and propose GRAPHCLEANER, a post-hoc method to detect and correct these mislabelled nodes in graph datasets. GRAPHCLEANER combines the novel ideas of 1) Synthetic Mislabel Dataset Generation, which seeks to generate realistic mislabels; and 2) Neighborhood-Aware Mislabel Detection, where neighborhood dependency is exploited in both labels and base classifier predictions. Empirical evaluations on 6 datasets and 6 experimental settings demonstrate that GRAPH-CLEANER outperforms the closest baseline, with an average improvement of 0.14 in F1 score, and 0.16 in MCC. On real-data case studies, GRAPH-CLEANER detects real and previously unknown mislabels in popular graph benchmarks: PubMed, Cora, CiteSeer and OGB-arxiv; we find that at least 6.91% of PubMed data is mislabelled or ambiguous, and simply removing these mislabelled data can boost evaluation performance from 86.71% to 89.11% 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Identifying Mislabeled Data using the Area Under the Margin RankingGeoff Pleiss, Tianyi Zhang, Ethan R. Elenberg, Kilian Q. WeinbergerNeurIPS 2020 · 398 citations
- Combining Label Propagation and Simple Models out-performs Graph Neural NetworksQian Huang, Horace He, Abhay Singh, Ser-Nam Lim et al.ICLR 2021 · 322 citations
- Be Confident! Towards Trustworthy Graph Neural Networks via Confidence CalibrationXiao Wang, Hongrui Liu, Chuan Shi, Cheng YangNeurIPS 2021 · 158 citations
- Clusterability as an Alternative to Anchor Points When Learning with Noisy LabelsZhaowei Zhu, Yiwen Song, Yang LiuICML 2021 · 112 citations
Related papers
- Resurrecting Label Propagation for Graphs with Heterophily and Label NoiseYao Cheng, Caihua Shan, Yifei Shen, Xiang Li et al.KDD 2024 · 8 citations
- Generated Graph DetectionYihan Ma, Zhikun Zhang, Ning Yu, Xinlei He et al.ICML 2023
- GALE: Active Adversarial Learning for Erroneous Node Detection in GraphsSheng Guan, Hanchao Ma, Mengying Wang, Yinghui WuICDE 2023 · 2 citations
- Neural Relation Graph: A Unified Framework for Identifying Label Noise and Outlier DataJang-Hyun Kim, Sangdoo Yun, Hyun Oh SongNeurIPS 2023 · 32 citations
- NGC: A Unified Framework for Learning with Open-World Noisy DataZhi-Fan Wu, Tong Wei, Jianwen Jiang, Chaojie Mao et al.ICCV 2021 · 106 citations
