Bridging Vision and Language Encoders: Parameter-Efficient Tuning for Referring Image Segmentation
Zunnan Xu, Zhihong Chen, Yong Zhang, Yibing Song, Xiang Wan, Guanbin Li
Abstract
Parameter Efficient Tuning (PET) has gained attention for reducing the number of parameters while maintaining performance and providing better hardware resource savings, but few studies investigate dense prediction tasks and interaction between modalities. In this paper, we do an investigation of efficient tuning problems on referring image segmentation. We propose a novel adapter called Bridger to facilitate cross-modal information exchange and inject task-specific information into the pre-trained model. We also design a lightweight decoder for image segmentation. Our approach achieves comparable or superior performance with only 1.61% to 3.38% backbone parameter updates, evaluated on challenging benchmarks. The code is available at https://github.com/kkakkkka/ETRIS .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db962256-e669-462a-9664-6ffb79f9f104Cited by top-tier papers23
- MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space ModelsZunnan Xu, Yukang Lin, Haonan Han, Sicheng Yang et al.NeurIPS 2024 · 62 citations
- Real-world Image Dehazing with Coherence-based Pseudo Labeling and Cooperative Unfolding NetworkChengyu Fang, Chunming He, Fengyang Xiao, Yulun Zhang et al.NeurIPS 2024 · 46 citations
- SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement LearningJiaqi Huang, Zunnan Xu, Jun Zhou, Ting Liu et al.NeurIPS 2025 · 33 citations
- Consistent123: One Image to Highly Consistent 3D Asset Using Case-Aware Diffusion PriorsYukang Lin, Haonan Han, Chaoqun Gong, Zunnan Xu et al.ACM MM 2024 · 20 citations
- Chain of Generation: Multi-Modal Gesture Synthesis via Cascaded Conditional ControlZunnan Xu, Yachao Zhang, Sicheng Yang, Ronghui Li et al.AAAI 2024 · 20 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang et al.NeurIPS 2022 · 1,291 citations
Related papers
- BarLeRIa: An Efficient Tuning Framework for Referring Image SegmentationYaoming Wang, Jin Li, Xiaopeng Zhang, Bowen Shi et al.ICLR 2024 · 11 citations
- Densely Connected Parameter-Efficient Tuning for Referring Image SegmentationJiaqi Huang, Zunnan Xu, Ting Liu, Yong Liu et al.AAAI 2025 · 34 citations
- VL-PET: Vision-and-Language Parameter-Efficient Tuning via Granularity ControlZi-Yuan Hu, Yanyang Li, Michael R. Lyu, Liwei WangICCV 2023 · 25 citations
- Dynamic Tuning Towards Parameter and Inference Efficiency for ViT AdaptationWangbo Zhao, Jiasheng Tang, Yizeng Han, Yibing Song et al.NeurIPS 2024 · 41 citations
- Faster Parameter-Efficient Tuning with Token Redundancy ReductionKwonyoung Kim, Jungin Park, Jin Kim, Hyeongjun Kwon et al.CVPR 2025
