No Metric to Rule Them All: Toward Principled Evaluations of Graph-Learning Datasets
Corinna Coupette, Jeremy Wayland, Emily Simons, Bastian Rieck
Abstract
Benchmark datasets have proved pivotal to the success of graph learning, and good benchmark datasets are crucial to guide the development of the field. Recent research has highlighted problems with graph-learning datasets and benchmarking practices-revealing, for example, that methods which ignore the graph structure can outperform graph-based approaches. Such findings raise two questions: (1) What makes a good graphlearning dataset, and (2) how can we evaluate dataset quality in graph learning? Our work addresses these questions. As the classic evaluation setup uses datasets to evaluate models, it does not apply to dataset evaluation. Hence, we start from first principles. Observing that graph-learning datasets uniquely combine two modes-graph structure and node features-, we introduce RINGS, a flexible and extensible modeperturbation framework to assess the quality of graph-learning datasets based on dataset ablations-i.e., quantifying differences between the original dataset and its perturbed representations. Within this framework, we propose two measures-performance separability and mode complementarity-as evaluation tools, each assessing the capacity of a graph dataset to benchmark the power and efficacy of graph-learning methods from a distinct angle. We demonstrate the utility of our framework for dataset evaluation via extensive experiments on graph-level tasks and derive actionable recommendations for improving the evaluation of graph-learning methods. Our work opens new research directions in datacentric graph learning, and it constitutes a step toward the systematic evaluation of evaluations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Geometry-Aware Edge Pooling for Graph Neural NetworksKatharina Limbeck, Lydia Mezrag, Guy Wolf, Bastian RieckNeurIPS 2025 · 9 citations
- Fixed Aggregation Features Can Rival GNNsCelia Rubio-Madrigal, Rebekka BurkholzICML 2026 · 2 citations
- Rethinking GNNs and Missing Features: Challenges, Evaluation and a Robust SolutionFrancesco Ferrini, Veronica Lachi, Antonio Longa, Bruno Lepri et al.ICML 2026 · 1 citation
- Plain Transformers are Surprisingly Powerful Link PredictorsQuang Truong, Yu Song, Donald Loveland, Mingxuan Ju et al.ICML 2026
- Is Graph Mixup Beneficial? Investigating Interpolation And Empirical Performance of Graph Mixup MethodsSimon Forbat, Rainer GemullaICML 2026
Builds on20
- Recipe for a General, Powerful, Scalable Graph TransformerLadislav Rampásek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu et al.NeurIPS 2022 · 1,216 citations
- Revisiting Heterophily For Graph Neural NetworksSitao Luan, Chenqing Hua, Qincheng Lu, Jiaqi Zhu et al.NeurIPS 2022 · 351 citations
- GraphFormers: GNN-nested Transformers for Representation Learning on Textual GraphJunhan Yang, Zheng Liu, Shitao Xiao, Chaozhuo Li et al.NeurIPS 2021 · 262 citations
- On Over-Squashing in Message Passing Neural Networks: The Impact of Width, Depth, and TopologyFrancesco Di Giovanni, Lorenzo Giusti, Federico Barbero, Giulia Luise et al.ICML 2023 · 190 citations
- When Do Graph Neural Networks Help with Node Classification? Investigating the Homophily Principle on Node DistinguishabilitySitao Luan, Chenqing Hua, Minkai Xu, Qincheng Lu et al.NeurIPS 2023 · 118 citations
Related papers
- Evaluating Robustness and Uncertainty of Graph Models Under Structural Distributional ShiftsGleb Bazhenov, Denis Kuznedelev, Andrey Malinin, Artem Babenko et al.NeurIPS 2023 · 12 citations
- Centrality-guided Pre-training for GraphBin Liang, Shiwei Chen, Lin Gui, Hui Wang et al.ICLR 2025
- Analyzing Data-Centric Properties for Graph Contrastive LearningPuja Trivedi, Ekdeep Singh Lubana, Mark Heimann, Danai Koutra et al.NeurIPS 2022 · 13 citations
- A Metadata-Driven Approach to Understand Graph Neural NetworksTing Wei Li, Qiaozhu Mei, Jiaqi MaNeurIPS 2023 · 13 citations
- Evaluation Metrics for Graph Generative Models: Problems, Pitfalls, and Practical SolutionsLeslie O'Bray, Max Horn, Bastian Rieck, Karsten M. BorgwardtICLR 2022 · 51 citations
