Vision HGNN: An Image is More than a Graph of Nodes
Yan Han, Peihao Wang, Souvik Kundu, Ying Ding, Zhangyang Wang
Abstract
The realm of graph-based modeling has proven its adaptability across diverse real-world data types. However, its applicability to general computer vision tasks had been limited until the introduction of the Vision Graph Neural Network (ViG). ViG divides input images into patches, conceptualized as nodes, constructing a graph through connections to nearest neighbors. Nonetheless, this method of graph construction confines itself to simple pairwise relationships, leading to surplus edges and unwarranted memory and computation expenses. In this paper, we enhance ViG by transcending conventional "pairwise" linkages and harnessing the power of the hypergraph to encapsulate image information. Our objective is to encompass more intricate inter-patch associations. In both training and inference phases, we adeptly establish and update the hypergraph structure using the Fuzzy C-Means method, ensuring minimal computational burden. This augmentation yields the Vision HyperGraph Neural Network (ViHGNN). The model's efficacy is empirically substantiated through its state-of-the-art performance on both image classification and object detection tasks, courtesy of the hypergraph structure learning module that uncovers higher-order relationships. Our code is available at: https://github . com/VITA-Group/ViHGNN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2aeda5c7-6d71-4c0b-aa1c-e234a0fb7f3bCited by top-tier papers14
- GreedyViG: Dynamic Axial Graph Construction for Efficient Vision GNNsMustafa Munir, William Avery, Md Mostafijur Rahman, Radu MarculescuCVPR 2024 · 27 citations
- Efficiency Follows Global-Local DecouplingZhenyu Yang, Gensheng Pei, Tao Chen, Yichao Zhou et al.CVPR 2026 · 3 citations
- HyperPotter: Spell the Charm of High-Order Interactions in Audio Deepfake DetectionQing Wen, Haohao Li, Zhongjie Ba, Peng Cheng et al.ICML 2026 · 1 citation
- Commonality in Few: Few-Shot Multimodal Anomaly Detection via Hypergraph-Enhanced MemoryYuxuan Lin, Hanjing Yan, Xuan Tong, Yang Chang et al.AAAI 2026 · 1 citation
- Adaptive Learned Image Compression with Graph Neural NetworksYunuo Chen, Bing He, Zezheng Lyu, Hongwei Hu et al.CVPR 2026 · 1 citation
Builds on31
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
Related papers
- DVHGNN: Multi-Scale Dilated Vision HGNN for Efficient Vision RecognitionCaoshuo Li, Tanzhe Li, Xiaobin Hu, Donghao Luo et al.CVPR 2025
- Vision GNN: An Image is Worth Graph of NodesKai Han, Yunhe Wang, Jianyuan Guo, Yehui Tang et al.NeurIPS 2022 · 668 citations
- Hypergraph Vision Transformers: Images are More than Nodes, More than EdgesJoshua FixelleCVPR 2025
- Self-Supervised Vision Graph Neural Networks Based on Contrastive LearningYuzhen Li, Yuehui Han, Jianjun Qian, Jian YangACM MM 2025
- AdaHGNN: Adaptive Hypergraph Neural Networks for Multi-Label Image ClassificationXiangping Wu, Qingcai Chen, Wei Li, Yulun Xiao et al.ACM MM 2020 · 53 citations
