LLMs Are Noisy Oracles! LLM-based Noise-aware Graph Active Learning for Node Classification
Zeang Sheng, Weiyang Guo, Yingxia Shao, Wentao Zhang, Bin Cui
Abstract
Graph Neural Networks (GNNs) have achieved significant success in various graph learning tasks. However, their success relies heavily on the availability of adequate high-quality labels, which requires high annotation costs. Recently, Large Language Models (LLMs) became known to the community for their strong zero-shot performance on diverse textual tasks. To utilize LLM's zero-shot abilities, existing work substitutes the oracle in the graph active learning setting with an LLM and achieves low-cost annotation compared to consulting human experts. However, annotations from LLMs inevitably contain noise, and we find through an empirical analysis that LLMs possess complex and dataset-specific annotation noise distributions. This intricacy of LLMs' annotation noise is neglected by existing work, leading to sub-optimal annotation data quality under the same budget. In this paper, we propose a Dataset- and LLM-Aware graph active learning framework called DMA that explicitly models LLM's annotation noise distributions on the given dataset. Concretely, DMA queries the oracle LLM with manually constructed prompts for each class's pseudo samples that are then used to measure the class-wise annotation noise rates. We then equip DMA with a new noise-aware graph active learning algorithm to utilize this fine-grained annotation noise distribution. Extensive performance optimizations are applied to the implementation to boost DMA's scalability and runtime efficiency on large-scale datasets. Evaluations on five text-attributed graph datasets show that DMA consistently outperforms all the baselines in terms of annotation data quality, which illustrates the importance of careful handling of LLM's intricate annotation noise.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get cf08f166-8f1c-4b4f-b6c6-fee8069e14eeCited by top-tier papers1
Ask how each one uses itRelated papers
- Leveraging Large Language Models for Effective Label-free Node Classification in Text-Attributed GraphsTaiyan Zhang, Renchi Yang, Yurui Lai, Mingyu Yan et al.SIGIR 2025 · 3 citations
- Test-Time Training on Graphs with Large Language Models (LLMs)Jiaxin Zhang, Yiqi Wang, Xihong Yang, Siwei Wang et al.ACM MM 2024 · 6 citations
- Label-free Node Classification on Graphs with Large Language Models (LLMs)Zhikai Chen, Haitao Mao, Hongzhi Wen, Haoyu Han et al.ICLR 2024 · 103 citations
- Dynamic Bundling with Large Language Models for Zero-Shot Inference on Text-Attributed GraphsYusheng Zhao, Qixin Zhang, Xiao Luo, Weizhi Zhang et al.NeurIPS 2025 · 4 citations
- Next Generation Active Learning: Mixture of LLMs in the LoopYuanyuan Qi, Xiaohao Yang, Jueqing Lu, Guoxiang Guo et al.AAAI 2026
