A critical look at the evaluation of GNNs under heterophily: Are we really making progress?
Oleg Platonov, Denis Kuznedelev, Michael Diskin, Artem Babenko, Liudmila Prokhorenkova
摘要
Node classification is a classical graph machine learning task on which Graph Neural Networks (GNNs) have recently achieved strong results. However, it is often believed that standard GNNs only work well for homophilous graphs, i.e., graphs where edges tend to connect nodes of the same class. Graphs without this property are called heterophilous, and it is typically assumed that specialized methods are required to achieve strong performance on such graphs. In this work, we challenge this assumption. First, we show that the standard datasets used for evaluating heterophily-specific models have serious drawbacks, making results obtained by using them unreliable. The most significant of these drawbacks is the presence of a large number of duplicate nodes in the datasets squirrel and chameleon, which leads to train-test data leakage. We show that removing duplicate nodes strongly affects GNN performance on these datasets. Then, we propose a set of heterophilous graphs of varying properties that we believe can serve as a better benchmark for evaluating the performance of GNNs under heterophily. We show that standard GNNs achieve strong results on these heterophilous graphs, almost always outperforming specialized models. Our datasets and the code for reproducing our experiments are available at https://github.com/yandex-research/heterophilous-graphs . ISSUES WITH POPULAR HETEROPHILOUS DATASETS In this section, we revisit datasets commonly used for heterophilous node classification. As discussed in Section 2, the following six datasets are the most popular: Wikipedia networks squirrel and chameleon, actor co-occurrence in Wikipedia pages network (actor), and WebKB datasets texas, wisconsin, and cornell. The standard preprocessing of these datasets is done by Pei et al. (2020) . First, we note that these datasets only come from three sources; thus, they
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper160
- Simplifying and Empowering Transformers for Large-Graph RepresentationsQitian Wu, Wentao Zhao, Chenxiao Yang, Hengrui Zhang 等NeurIPS 2023 · 被引用 318 次
- Characterizing Graph Datasets for Node Classification: Homophily-Heterophily Dichotomy and BeyondOleg Platonov, Denis Kuznedelev, Artem Babenko, Liudmila ProkhorenkovaNeurIPS 2023 · 被引用 95 次
- Simple and Asymmetric Graph Contrastive Learning without AugmentationsTeng Xiao, Huaisheng Zhu, Zhengyu Chen, Suhang WangNeurIPS 2023 · 被引用 86 次
- Polynormer: Polynomial-Expressive Graph Transformer in Linear TimeChenhui Deng, Zichao Yue, Zhiru ZhangICLR 2024 · 被引用 81 次
- A Fractional Graph Laplacian Approach to OversmoothingSohir Maskey, Raffaele Paolino, Aras Bacho, Gitta KutyniokNeurIPS 2023 · 被引用 66 次
它引用的顶会 Paper17
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- How Attentive are Graph Attention Networks?Shaked Brody, Uri Alon, Eran YahavICLR 2022 · 被引用 1,717 次
- Beyond Homophily in Graph Neural Networks: Current Limitations and Effective DesignsJiong Zhu, Yujun Yan, Lingxiao Zhao, Mark Heimann 等NeurIPS 2020 · 被引用 1,490 次
- Geom-GCN: Geometric Graph Convolutional NetworksHongbin Pei, Bingzhe Wei, Kevin Chen-Chuan Chang, Yu Lei 等ICLR 2020 · 被引用 1,445 次
- Beyond Low-frequency Information in Graph Convolutional NetworksDeyu Bo, Xiao Wang, Chuan Shi, Huawei ShenAAAI 2021 · 被引用 773 次
相关 Paper
- Is Homophily a Necessity for Graph Neural Networks?Yao Ma, Xiaorui Liu, Neil Shah, Jiliang TangICLR 2022 · 被引用 295 次
- On the Impact of Feature Heterophily on Link Prediction with Graph Neural NetworksJiong Zhu, Gaotang Li, Yao-An Yang, Jing Zhu 等NeurIPS 2024 · 被引用 21 次
- AGS-GNN: Attribute-guided Sampling for Graph Neural NetworksSiddhartha Shankar Das, S. M. Ferdous, Mahantesh M. Halappanavar, Edoardo Serra 等KDD 2024 · 被引用 3 次
- Large Scale Learning on Non-Homophilous Graphs: New Benchmarks and Strong Simple MethodsDerek Lim, Felix Hohne, Xiuyu Li, Sijia Linda Huang 等NeurIPS 2021 · 被引用 534 次
- Revisiting Heterophily For Graph Neural NetworksSitao Luan, Chenqing Hua, Qincheng Lu, Jiaqi Zhu 等NeurIPS 2022 · 被引用 351 次
