Vision Graph Prompting via Semantic Low-Rank Decomposition
Zixiang Ai, Zichen Liu, Jiahuan Zhou
Abstract
Vision GNN (ViG) demonstrates superior performance by representing images as graph structures, providing a more natural way to capture irregular semantic patterns beyond traditional grid or sequence-based representations. To efficiently adapt ViG to downstream tasks, parameterefficient fine-tuning techniques like visual prompting become increasingly essential. However, existing prompting methods are primarily designed for Transformer-based models, neglecting the rich topological relationships among nodes and edges in graph-based representations, limiting their capacity to model complex semantics. In this paper, we propose Vision Graph Prompting (VGP), a novel framework tailored for vision graph structures. Our core insight reveals that semantically connected components in the graph exhibit low-rank properties. Building on this observation, we introduce a semantic low-rank prompting method that decomposes low-rank semantic features and integrates them with prompts on vision graph topologies, capturing both global structural patterns and fine-grained semantic dependencies. Extensive experiments demonstrate our method significantly improves ViG's transfer performance on diverse downstream tasks, achieving results comparable to full fine-tuning while maintaining parameter efficiency. Our code is available at https://github.com/ zhoujiahuan1991/ICML2025-VGP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e95ab5b2-6ae7-452d-8601-04a711456ef7Cited by top-tier papers1
Ask how each one uses itBuilds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik et al.ICLR 2020 · 1,744 citations
- DeepGCNs: Can GCNs Go As Deep As CNNs?Guohao Li, Matthias Müller, Ali K. Thabet, Bernard GhanemICCV 2019 · 1,586 citations
Related papers
- Vision GNN: An Image is Worth Graph of NodesKai Han, Yunhe Wang, Jianyuan Guo, Yehui Tang et al.NeurIPS 2022 · 668 citations
- AutoVP: An Automated Visual Prompting Framework and BenchmarkHsi-Ai Tsao, Lei Hsiung, Pin-Yu Chen, Si Liu et al.ICLR 2024 · 29 citations
- Selective Visual Prompting in Vision MambaYifeng Yao, Zichen Liu, Zhenyu Cui, Yuxin Peng et al.AAAI 2025 · 13 citations
- LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model AdaptationCan Jin, Ying Li, Mingyu Zhao, Shiyu Zhao et al.ICLR 2025
- Visual Instance-aware Prompt TuningXi Xiao, Yunbei Zhang, Xingjian Li, Tianyang Wang et al.ACM MM 2025 · 12 citations
