LLMs Are Noisy Oracles! LLM-based Noise-aware Graph Active Learning for Node Classification
Zeang Sheng, Weiyang Guo, Yingxia Shao, Wentao Zhang, Bin Cui
摘要
Graph Neural Networks (GNNs) have achieved significant success in various graph learning tasks. However, their success relies heavily on the availability of adequate high-quality labels, which requires high annotation costs. Recently, Large Language Models (LLMs) became known to the community for their strong zero-shot performance on diverse textual tasks. To utilize LLM's zero-shot abilities, existing work substitutes the oracle in the graph active learning setting with an LLM and achieves low-cost annotation compared to consulting human experts. However, annotations from LLMs inevitably contain noise, and we find through an empirical analysis that LLMs possess complex and dataset-specific annotation noise distributions. This intricacy of LLMs' annotation noise is neglected by existing work, leading to sub-optimal annotation data quality under the same budget. In this paper, we propose a Dataset- and LLM-Aware graph active learning framework called DMA that explicitly models LLM's annotation noise distributions on the given dataset. Concretely, DMA queries the oracle LLM with manually constructed prompts for each class's pseudo samples that are then used to measure the class-wise annotation noise rates. We then equip DMA with a new noise-aware graph active learning algorithm to utilize this fine-grained annotation noise distribution. Extensive performance optimizations are applied to the implementation to boost DMA's scalability and runtime efficiency on large-scale datasets. Evaluations on five text-attributed graph datasets show that DMA consistently outperforms all the baselines in terms of annotation data quality, which illustrates the importance of careful handling of LLM's intricate annotation noise.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Leveraging Large Language Models for Effective Label-free Node Classification in Text-Attributed GraphsTaiyan Zhang, Renchi Yang, Yurui Lai, Mingyu Yan 等SIGIR 2025 · 被引用 3 次
- Test-Time Training on Graphs with Large Language Models (LLMs)Jiaxin Zhang, Yiqi Wang, Xihong Yang, Siwei Wang 等ACM MM 2024 · 被引用 6 次
- Label-free Node Classification on Graphs with Large Language Models (LLMs)Zhikai Chen, Haitao Mao, Hongzhi Wen, Haoyu Han 等ICLR 2024 · 被引用 103 次
- Dynamic Bundling with Large Language Models for Zero-Shot Inference on Text-Attributed GraphsYusheng Zhao, Qixin Zhang, Xiao Luo, Weizhi Zhang 等NeurIPS 2025 · 被引用 4 次
- Next Generation Active Learning: Mixture of LLMs in the LoopYuanyuan Qi, Xiaohao Yang, Jueqing Lu, Guoxiang Guo 等AAAI 2026
