Bonsai: Gradient-free Graph Condensation for Node Classification
Mridul Gupta, Samyak Jain, Vansh Ramani, Hariprasad Kodamana, Sayan Ranu
Abstract
Graph condensation has emerged as a promising avenue to enable scalable training of Gnns by compressing the training dataset while preserving essential graph characteristics. Our study uncovers significant shortcomings in current graph condensation techniques. First, the majority of the algorithms paradoxically require training on the full dataset to perform condensation. Second, due to their gradient-emulating approach, these methods require fresh condensation for any change in hyper-parameters or Gnn architecture, limiting their flexibility and reusability. To address these challenges, we present Bonsai, a novel graph condensation method empowered by the observation that computation trees form the fundamental processing units of message-passing Gnns. Bonsai condenses datasets by encoding a careful selection of exemplar trees that maximize the representation of all computation trees in the training set. This unique approach imparts Bonsai as the first linear-time, model-agnostic graph condensation algorithm for node classification that outperforms existing baselines across 7 real-world datasets on accuracy, while being 22 times faster on average. Bonsai is grounded in rigorous mathematical guarantees on the adopted approximation strategies, making it robust to Gnn architectures, datasets, and parameters. *Denotes equal contribution. ¹Some algorithms sparsify the fully-connected graph based on edge weights. But this sparsification process requires training on the fully connected graph itself to identify the pruning threshold. ²Inspired by the art of Bonsai, which transforms large trees into miniature forms while preserving their essence, our graph condensation algorithm gracefully prunes redundant computation trees, creating a condensed graph that is significantly smaller yet maintains comparable performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on15
- GraphSAINT: Graph Sampling Based Inductive Learning MethodHanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava, Rajgopal Kannan et al.ICLR 2020 · 1,155 citations
- Graph Condensation for Graph Neural NetworksWei Jin, Lingxiao Zhao, Shichang Zhang, Yozen Liu et al.ICLR 2022 · 203 citations
- Condensing Graphs via One-Step Gradient MatchingWei Jin, Xianfeng Tang, Haoming Jiang, Zheng Li et al.KDD 2022 · 68 citations
- Does Graph Distillation See Like Vision Dataset Counterpart?Beining Yang, Kai Wang, Qingyun Sun, Cheng Ji et al.NeurIPS 2023 · 62 citations
- Fast Graph Condensation with Structure-based Neural Tangent KernelLin Wang, Wenqi Fan, Jiatong Li, Yao Ma et al.WWW 2024 · 45 citations
Related papers
- Mirage: Model-agnostic Graph Distillation for Graph ClassificationMridul Gupta, Sahil Manchanda, Hariprasad Kodamana, Sayan RanuICLR 2024 · 17 citations
- Bi-Directional Multi-Scale Graph Dataset Condensation via Information BottleneckXingcheng Fu, Yisen Gao, Beining Yang, Yuxuan Wu et al.AAAI 2025 · 7 citations
- Disentangled Condensation for Large-scale GraphsZhenbang Xiao, Yu Wang, Shunyu Liu, Bingde Hu et al.WWW 2025 · 14 citations
- Graph Condensation for Inductive Node Representation LearningXinyi Gao, Tong Chen, Yilong Zang, Wentao Zhang et al.ICDE 2024 · 32 citations
- Rethinking and Accelerating Graph Condensation: A Training-Free Approach with Class PartitionXinyi Gao, Guanhua Ye, Tong Chen, Wentao Zhang et al.WWW 2025 · 27 citations
