Progressive Parameter Efficient Transfer Learning for Semantic Segmentation
Nan Zhou, Huiqun Wang, Yaoyan Zheng, Di Huang
Abstract
Parameter Efficient Transfer Learning (PETL) excels in downstream classification fine-tuning with minimal computational overhead, demonstrating its potential within the pre-train and fine-tune paradigm. However, recent PETL methods consistently struggle when fine-tuning for semantic segmentation tasks, limiting their broader applicability. In this paper, we identify that fine-tuning for semantic segmentation requires larger parameter adjustments due to shifts in semantic perception granularity. Current PETL approaches are unable to effectively accommodate these shifts, leading to significant performance degradation. To address this, we introduce ProPETL, a novel approach that incorporates an additional midstream adaptation to progressively align pre-trained models for segmentation tasks. Through this process, ProPETL achieves state-of-the-art performance on most segmentation benchmarks and, for the first time, surpasses full fine-tuning on the challenging COCO-Stuff10k dataset. Furthermore, ProPETL demonstrates strong generalization across various pre-trained models and scenarios, highlighting its effectiveness and versatility for broader adoption in segmentation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 78dd6950-b6de-4e6b-ae49-87c7aabf6b41Cited by top-tier papers3
- Implicit Modeling for Transferability Estimation of Vision Foundation ModelsYaoyan Zheng, Huiqun Wang, Nan Zhou, Di HuangNeurIPS 2025 · 1 citation
- What Makes Synthetic Data Effective in Image SegmentationJinjin Zhang, Xiefan Guo, Yizhou jin, Nan Zhou et al.ICML 2026
- CoVFT: Context-aware Visual Fine-tuning for Multimodal Large Language ModelsNan Zhou, Huiqun Wang, Yaoyan Zheng, Di HuangCVPR 2026
Builds on14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
Related papers
- Parameter-efficient is not Sufficient: Exploring Parameter, Memory, and Time Efficient Adapter Tuning for Dense PredictionsDongshuo Yin, Xueting Han, Bin Li, Hao Feng et al.ACM MM 2024 · 18 citations
- Promptable Anomaly Segmentation with SAM Through Self-Perception TuningHui-Yue Yang, Hui Chen, Ao Wang, Kai Chen et al.AAAI 2025 · 10 citations
- Exploring Vision Semantic Prompt for Efficient Point Cloud UnderstandingYixin Zha, Chuxin Wang, Wenfei Yang, Tianzhu Zhang et al.ICML 2025
- One Network, Many Masks: Towards More Parameter-Efficient Transfer LearningGuangtao Zeng, Peiyuan Zhang, Wei LuACL 2023 · 8 citations
- Semantically-Shifted Incremental Adapter-Tuning is A Continual ViTransformerYuwen Tan, Qinhao Zhou, Xiang Xiang, Ke Wang et al.CVPR 2024 · 14 citations
