Cost-effective Data Labelling for Graph Neural Networks
Shixun Huang, Ge Lee, Zhifeng Bao, Shirui Pan
摘要
Active learning (AL), that aims to label limited data samples to effectively train the model, stands as a very cost-effective data labelling strategy in machine learning. Given the state-of-the-art performance GNNs have achieved in graph-based tasks, it is critical to design proper AL methods for graph neural networks (GNNs). However, existing GNN-based AL methods require considerable supervised information to guide the AL process, such as the GNN model to use, and initially labelled nodes and labels of newly selected nodes. Such dependency on supervised information limits both flexibility and scalabilty. In this paper, we propose an unsupervised, scalable and flexible AL method - it incurs low memory footprints and time cost, is flexible to the choice of underlying GNNs, and operates without requiring GNN-model-specific knowledge or labels of selected nodes. Specifically, we leverage the commonality of existing GNNs to reformulate the unsupervised AL problem as the Aggregation Involvement Maximization (AIM) problem. The objective of AIM is to maximize the involvement or participation of all nodes during the feature aggregation process of GNNs for nodes to be labelled. In this way, the aggregated features of labelled nodes can be diversified to a large extent, thereby benefiting the training of feature transformation matrices which are major trainable components in GNNs. We prove that the AIM problem is NP-hard and propose an efficient solution with theoretical guarantees. Extensive experiments on public datasets demonstrate the effectiveness, scalability and flexibility of our method. Our study is highly relevant to the track "Graph Algorithms and Modeling for the Web" since we focus one of the major listed topics "Graph Embedding and GNNs for the Web" and AL for GNNs, as an important research problem, is faced by aforementioned challenges to be tackled in this paper.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- No Change, No Gain: Empowering Graph Neural Networks with Expected Model Change Maximization for Active LearningZixing Song, Yifei Zhang, Irwin KingNeurIPS 2023 · 被引用 21 次
- Information Gain Propagation: a New Way to Graph Active Learning with Soft LabelsWentao Zhang, Yexin Wang, Zhenbang You, Meng Cao 等ICLR 2022 · 被引用 24 次
- ALG: Fast and Accurate Active Learning Framework for Graph Convolutional NetworksWentao Zhang, Yu Shen, Yang Li, Lei Chen 等SIGMOD 2021 · 被引用 36 次
- NC-ALG: Graph-Based Active Learning Under Noisy CrowdWentao Zhang, Yexin Wang, Zhenbang You, Yang Li 等ICDE 2024 · 被引用 4 次
- Graph Policy Network for Transferable Active Learning on GraphsShengding Hu, Zheng Xiong, Meng Qu, Xingdi Yuan 等NeurIPS 2020 · 被引用 84 次
