Partition and Code: learning how to compress graphs
Giorgos Bouritsas, Andreas Loukas, Nikolaos Karalias, Michael M. Bronstein
Abstract
Can we use machine learning to compress graph data? The absence of ordering in graphs poses a significant challenge to conventional compression algorithms, limiting their attainable gains as well as their ability to discover relevant patterns. On the other hand, most graph compression approaches rely on domain-dependent handcrafted representations and cannot adapt to different underlying graph distributions. This work aims to establish the necessary principles a lossless graph compression method should follow to approach the entropy storage lower bound. Instead of making rigid assumptions about the graph distribution, we formulate the compressor as a probabilistic model that can be learned from data and generalise to unseen instances. Our "Partition and Code" framework entails three steps: first, a partitioning algorithm decomposes the graph into subgraphs, then these are mapped to the elements of a small dictionary on which we learn a probability distribution, and finally, an entropy encoder translates the representation into bits. All the components (partitioning, dictionary and distribution) are parametric and can be trained with gradient descent. We theoretically compare the compression quality of several graph encodings and prove, under mild conditions, that PnC achieves compression gains that grow either linearly or quadratically with the number of vertices. Empirically, PnC yields significant compression improvements on diverse real-world networks. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42403dd2-abdf-49d4-8716-136dd75f1e24Cited by top-tier papers6
- Equivariant Subgraph Aggregation NetworksBeatrice Bevilacqua, Fabrizio Frasca, Derek Lim, Balasubramaniam Srinivasan et al.ICLR 2022 · 217 citations
- Entropy Coding of Unordered Data StructuresJulius Kunze, Daniel Severo, Giulio Zani, Jan-Willem van de Meent et al.ICLR 2024 · 7 citations
- Practical Shuffle CodingJulius Kunze, Daniel Severo, Jan-Willem van de Meent, James TownsendNeurIPS 2024 · 2 citations
- One-Shot Compression of Large Edge-Exchangeable Graphs using Bits-Back CodingDaniel Severo, James Townsend, Ashish J. Khisti, Alireza MakhzaniICML 2023 · 2 citations
- Graph Generation with K2-treesYunhui Jang, Dongwoo Kim, Sungsoo AhnICLR 2024 · 1 citation
Builds on16
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- GraphAF: a Flow-based Autoregressive Model for Molecular Graph GenerationChence Shi, Minkai Xu, Zhaocheng Zhu, Weinan Zhang et al.ICLR 2020 · 532 citations
- Spectral Clustering with Graph Neural Networks for Graph PoolingFilippo Maria Bianchi, Daniele Grattarola, Cesare AlippiICML 2020 · 528 citations
- Can Graph Neural Networks Count Substructures?Zhengdao Chen, Lei Chen, Soledad Villar, Joan BrunaNeurIPS 2020 · 392 citations
- What graph neural networks cannot learn: depth vs widthAndreas LoukasICLR 2020 · 336 citations
Related papers
- Few-Shot Non-Parametric Learning with Deep Latent Variable ModelZhiying Jiang, Yiqin Dai, Ji Xin, Ming Li et al.NeurIPS 2022 · 6 citations
- Learning Graph Representation via Graph Entropy MaximizationZiheng Sun, Xudong Wang, Chris Ding, Jicong FanICML 2024 · 9 citations
- Data Compression as a Comprehensive Framework for Graph Drawing and Representation LearningClaudia Plant, Sonja Biedermann, Christian BöhmKDD 2020 · 6 citations
- Learning Better Lossless Compression Using Lossy CompressionFabian Mentzer, Luc Van Gool, Michael TschannenCVPR 2020
- StructComp: Substituting propagation with Structural Compression in Training Graph Contrastive LearningShengzhong Zhang, Wenjie Yang, Xinyuan Cao, Hongwei Zhang et al.ICLR 2024 · 6 citations
