Parameter Efficient Fine-Tuning via Cross Block Orchestration for Segment Anything Model
Zelin Peng, Zhengqin Xu, Zhilin Zeng, Lingxi Xie, Qi Tian, Wei Shen
Abstract
Parameter-efficient fine-tuning (PEFT) is an effective methodology to unleash the potential of large foundation models in novel scenarios with limited training data. In the computer vision community, PEFT has shown effectiveness in image classification, but little research has studied its ability for image segmentation. Fine-tuning segmentation models usually requires a heavier adjustment of parameters to align the proper projection directions in the parameter space for new scenarios. This raises a challenge to existing PEFT algorithms, as they often inject a limited number of individual parameters into each block, which prevents substantial adjustment of the projection direction of the parameter space due to the limitation of Hidden Markov Chain along blocks. In this paper, we equip PEFT with a cross-block orchestration mechanism to enable the adaptation of the Segment Anything Model (SAM) to various downstream scenarios. We introduce a novel inter-block communication module, which integrates a learnable relation matrix to facilitate communication among different coefficient sets of each PEFT block's parameter space. Moreover, we propose an intra-block enhancement module, which introduces a linear projection head whose weights are generated from a hyper-complex layer, further enhancing the impact of the adjustment of projection directions on the entire parameter space. Extensive experiments on diverse benchmarks demonstrate that our proposed approach consistently improves the segmentation performance significantly on novel scenarios with only around 1K additional parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2fd3accf-44b1-4918-b9c3-13af41dcea2cCited by top-tier papers11
- Segment Anything in 3D with NeRFsJiazhong Cen, Zanwei Zhou, Jiemin Fang, Chen Yang et al.NeurIPS 2023 · 255 citations
- Robust SAM: On the Adversarial Robustness of Vision Foundation ModelsJiahuan Long, Zhengqin Xu, Tingsong Jiang, Wen Yao et al.AAAI 2025 · 5 citations
- S2FT: Parameter-Efficient Fine-Tuning in Sparse Spectrum DomainBaoquan Zhang, Zhehao Yu, Lisai Zhang, Kenghong Lin et al.CVPR 2026 · 1 citation
- InfoSAM: Fine-Tuning the Segment Anything Model from An Information-Theoretic PerspectiveYuanhong Zhang, Muyao Yuan, Weizhan Zhang, Tieliang Gong et al.ICML 2025
- Intervening Anchor Token: Decoding Strategy in Alleviating Hallucinations for MLLMsFeilong Tang, Zile Huang, Chengzhi Liu, Qiang Sun et al.ICLR 2025
Builds on18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang et al.NeurIPS 2022 · 1,291 citations
Related papers
- Promptable Anomaly Segmentation with SAM Through Self-Perception TuningHui-Yue Yang, Hui Chen, Ao Wang, Kai Chen et al.AAAI 2025 · 10 citations
- SAM-PARSER: Fine-Tuning SAM Efficiently by Parameter Space ReconstructionZelin Peng, Zhengqin Xu, Zhilin Zeng, Xiaokang Yang et al.AAAI 2024 · 41 citations
- Parameter-Free Fine-tuning via Redundancy Elimination for Vision Foundation ModelsJiahuan Long, Tingsong Jiang, Wen Yao, Yizhe Xiong et al.AAAI 2026
- VFM-Adapter: Adapting Visual Foundation Models for Dense Prediction with Dynamic Hybrid Operation MappingZheng Chen, Yu Zeng, Zehui Chen, Hongzhi Gao et al.AAAI 2025 · 1 citation
- Uncertainty-aware Fine-tuning of Segmentation Foundation ModelsKangning Liu, Brian L. Price, Jason Kuen, Yifei Fan et al.NeurIPS 2024 · 15 citations
