GraphCleaner: Detecting Mislabelled Samples in Popular Graph Learning Benchmarks
Yuwen Li, Miao Xiong, Bryan Hooi
摘要
Label errors have been found to be prevalent in popular text, vision, and audio datasets, which heavily influence the safe development and evaluation of machine learning algorithms. Despite increasing efforts towards improving the quality of generic data types, such as images and texts, the problem of mislabel detection in graph data remains underexplored. To bridge the gap, we explore mislabelling issues in popular real-world graph datasets and propose GRAPHCLEANER, a post-hoc method to detect and correct these mislabelled nodes in graph datasets. GRAPHCLEANER combines the novel ideas of 1) Synthetic Mislabel Dataset Generation, which seeks to generate realistic mislabels; and 2) Neighborhood-Aware Mislabel Detection, where neighborhood dependency is exploited in both labels and base classifier predictions. Empirical evaluations on 6 datasets and 6 experimental settings demonstrate that GRAPH-CLEANER outperforms the closest baseline, with an average improvement of 0.14 in F1 score, and 0.16 in MCC. On real-data case studies, GRAPH-CLEANER detects real and previously unknown mislabels in popular graph benchmarks: PubMed, Cora, CiteSeer and OGB-arxiv; we find that at least 6.91% of PubMed data is mislabelled or ambiguous, and simply removing these mislabelled data can boost evaluation performance from 86.71% to 89.11% 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- Identifying Mislabeled Data using the Area Under the Margin RankingGeoff Pleiss, Tianyi Zhang, Ethan R. Elenberg, Kilian Q. WeinbergerNeurIPS 2020 · 被引用 398 次
- Combining Label Propagation and Simple Models out-performs Graph Neural NetworksQian Huang, Horace He, Abhay Singh, Ser-Nam Lim 等ICLR 2021 · 被引用 322 次
- Be Confident! Towards Trustworthy Graph Neural Networks via Confidence CalibrationXiao Wang, Hongrui Liu, Chuan Shi, Cheng YangNeurIPS 2021 · 被引用 158 次
- Clusterability as an Alternative to Anchor Points When Learning with Noisy LabelsZhaowei Zhu, Yiwen Song, Yang LiuICML 2021 · 被引用 112 次
相关 Paper
- Resurrecting Label Propagation for Graphs with Heterophily and Label NoiseYao Cheng, Caihua Shan, Yifei Shen, Xiang Li 等KDD 2024 · 被引用 8 次
- Generated Graph DetectionYihan Ma, Zhikun Zhang, Ning Yu, Xinlei He 等ICML 2023
- GALE: Active Adversarial Learning for Erroneous Node Detection in GraphsSheng Guan, Hanchao Ma, Mengying Wang, Yinghui WuICDE 2023 · 被引用 2 次
- Neural Relation Graph: A Unified Framework for Identifying Label Noise and Outlier DataJang-Hyun Kim, Sangdoo Yun, Hyun Oh SongNeurIPS 2023 · 被引用 32 次
- NGC: A Unified Framework for Learning with Open-World Noisy DataZhi-Fan Wu, Tong Wei, Jianwen Jiang, Chaojie Mao 等ICCV 2021 · 被引用 106 次
