GraphMaster: Automated Graph Synthesis via LLM Agents in Data-Limited Environments
Enjun Du, Xunkai Li, Tian Jin, Zhihan Zhang, Rong-Hua Li, Guoren Wang
Abstract
The era of foundation models has revolutionized AI research, yet Graph Foundation Models (GFMs) remain constrained by the scarcity of large-scale graph corpora. Traditional graph data synthesis techniques primarily focus on simplistic structural operations, lacking the capacity to generate semantically rich nodes with meaningful textual attributes-a critical limitation for real-world applications. While large language models (LLMs) demonstrate exceptional text generation capabilities, their direct application to graph synthesis is impeded by context window limitations, hallucination phenomena, and structural consistency challenges. To address these issues, we introduce GraphMaster-the first multi-agent framework specifically designed for graph data synthesis in data-limited environments. GraphMaster orchestrates four specialized LLM agents (Manager, Perception, Enhancement, and Evaluation) that collaboratively optimize the synthesis process through iterative refinement, ensuring both semantic coherence and structural integrity. To rigorously evaluate our approach, we create new data-limited "Sub" variants of six standard graph benchmarks, specifically designed to test synthesis capabilities under realistic constraints. Additionally, we develop a novel interpretability assessment framework that combines human evaluation with a principled Grassmannian manifold-based analysis, providing both qualitative and quantitative measures of semantic coherence. Experimental results demonstrate that GraphMaster significantly outperforms traditional synthesis methods across multiple datasets, establishing a strong foundation for advancing GFMs in data-scarce environments. 2 * Corresponding author 2 Code is available on https://github.com/EnjunDu/GraphMaster. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
existing connections but cannot create novel nodes or patterns. Node-level mixing techniques like GraphMixup [53] generate synthetic nodes by interpolating features but often produce semantically inconsistent attributes, particularly with textual features. Graph-level synthesis methods such as G-Mixup [18] create entirely new graphs but struggle to balance global structure with local semantic coherence. The core limitation across these traditional methods is their inability to simultaneously preserve meaningful semantics while generating structurally valid expansions-a deficiency particularly pronounced when handling text-attributed graphs (TAGs) where both connectivity patterns and textual node features must remain coherent.
Large language models have demonstrated remarkable capabilities in understanding and generating text [19,11,61,24,2,45], suggesting potential for synthesizing text-attributed graphs. However, directly applying LLMs encounters several critical challenges: standard context windows cannot process entire graphs with numerous textual nodes [3]; LLMs excel at semantic understanding but struggle to maintain structural consistency [12]; and without proper coordination, they tend to produce inconsistent or hallucinated content that fails to capture the intricate balance between topology and semantics [33]. Furthermore, in realistic scenarios with limited available data, LLMs have insufficient examples to learn complex graph patterns [27,44,25].
To address these challenges, we propose GraphMaster, a novel multi-agent framework specifically designed for graph synthesis in data-limited environments. GraphMaster decomposes the complex synthesis task into specialized sub-tasks handled by four collaborative LLM-powered agents, each targeting specific challenges: (1) The Manager Agent coordinates the overall process and determines optimal synthesis strategies based on current graph characteristics, orchestrating the complex synthesis workflow; (2) The Perception Agent analyzes graph structure and employs advanced sampling to identify representative subgraphs processable within LLM context constraints, directly addressing the context window limitations; (3) The Enhancement Agent generates new nodes and edges with consistent semantics and structure, mitigating hallucination by maintaining coherence with existing graph elements; and (4) The Evaluation Agent assesses quality based on both semantic coherence and structural integrity, providing feedback for iterative improvement to ensure structural and semantic consistency. This decomposition enables targeted solutions for each challenge that a single-pass LLM approach cannot address.
Through this collaborative, iterative process, these specialized agents overcome the limitations of both traditional methods and direct LLM applications. The multi-agent architecture enables GraphMaster to effectively balance semantic richness with structural validity-producing high-quality synthetic graph data even with limited training examples. By introducing modular reasoning (through task decomposition), semantic control (via specialized agent expertise),
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b806fe2f-d83b-4ad6-b87f-a41635503147Cited by top-tier papers6
- GraphOracle: Efficient Fully-Inductive Knowledge Graph Reasoning via Relation-Dependency GraphsEnjun Du, Siyi Liu, Yongqi ZhangAAAI 2026 · 3 citations
- GRAPHIA: Harnessing Social Graph Data to Enhance LLM-Based Social SimulationJiarui Ji, Zehua Zhang, Zhewei Wei, Bin Tong et al.ACL 2026 · 1 citation
- Noise-Aware Graph-Based Cognitive Diagnostic Framework Through Low-Rank AlignmentGuixian Zhang, Yanmei Zhang, Guan Yuan, Shang Liu et al.AAAI 2026
- SelPE: Progressive Selection for Private Structured Text SynthesisXuancheng Zhu, Guoshun Nan, Han Zhang, Ben Niu et al.KDD 2026
- VideoSeg-R1: Reasoning Video Object Segmentation via Reinforcement LearningZishan Xu, Yifu Guo, Yuquan Lu, Fengyu Yang et al.AAAI 2026
Builds on24
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- Data Augmentation for Graph Neural NetworksTong Zhao, Yozen Liu, Leonardo Neves, Oliver J. Woodford et al.AAAI 2021 · 487 citations
- GPT-GNN: Generative Pre-Training of Graph Neural NetworksZiniu Hu, Yuxiao Dong, Kuansan Wang, Kai-Wei Chang et al.KDD 2020 · 438 citations
- G-Mixup: Graph Data Augmentation for Graph ClassificationXiaotian Han, Zhimeng Jiang, Ninghao Liu, Xia HuICML 2022 · 251 citations
- Knowledge Graph Reasoning with Relational DigraphYongqi Zhang, Quanming YaoWWW 2022 · 193 citations
Related papers
- GraphSynth: Resolving the Diversity-Reliability Trade-off with Probabilistic Factor GraphsZehua Cheng, Wei Dai, Jiahao Sun, Thomas LukasiewiczACL 2026
- LLaGA: Large Language and Graph AssistantRunjin Chen, Tong Zhao, Ajay Kumar Jaiswal, Neil Shah et al.ICML 2024 · 180 citations
- MA-GTS: A Multi-Agent Framework for Solving Complex Graph Problems in Real-World ApplicationsZike Yuan, Ming Liu, Hui Wang, Bing QinEMNLP 2025
- Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge GraphsYurun Chen, Xueyu Hu, Yuhan Liu, Ziqi Wang et al.CVPR 2026 · 7 citations
- Test-Time Search for Automated GFM Fine-TuningWenji Hu, Xianan Wang, Chunyu Wei, Senhao Liu et al.KDD 2026
