Unleashing the Power of Visual Foundation Models for Generalizable Semantic Segmentation
Peiyuan Tang, Xiaodong Zhang, Chunze Yang, Haoran Yuan, Jun Sun, Danfeng Shan, Zijiang James Yang
Abstract
Deep learning models often suffer from performance degradation in unseen domains, posing a risk for safety-critical applications such as autonomous driving. To tackle this problem, recent studies have leveraged pre-trained Visual Foundation Models (VFMs) to enhance generalization. However, exsiting works mainly focus on designing intricate networks for VFMs, neglecting their inherent strong generalization potential. Moreover, these methods typically perform inference on low-resolution images. The loss of detail hinders accurate predictions in unseen domains, especially for small objects. In this paper, we argue that simply fine-tuning VFMs and leveraging high-resolution images unleash the power of VFMs for generalizable semantic segmentation. Therefore, we design a VFM-based segmentation network (VFMNet) that adapts VFMs to this task with minimal fine-tuning, preserving their generalizable knowledge. Then, to fully utilize high-resolution images, we train a Mask-guided Refinement Network (MGRNet) to refine VFMNet's predictions combining detailed image features. Furthermore, we adopt a two-stage coarse-to-fine inference approach. MGRNet is used to refine the low-confidence regions predicted by VFMNet to obtain fine-grained results. Extensive experiments demonstrate the effectiveness of our method, outperforming state-of-the-art methods by 3.3% on the average mIoU in synthetic-to-real domain generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 877bdf3a-610a-415a-b867-211ef602bbe9Cited by top-tier papers2
- Learning a Cross-Modal Schrödinger Bridge for Visual Domain GeneralizationHao Zheng, Jingjun Yi, Qi Bi, Huimin Huang et al.NeurIPS 2025 · 1 citation
- Exploring Generalizable Remote Sensing Change Detection via Low-Rank Exchange Adaptation of Vision Foundation ModelMingwei Zhang, Jingtao Hu, Qiang Li, Qi WangAAAI 2026
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- DAFormer: Improving Network Architectures and Training Strategies for Domain-Adaptive Semantic SegmentationLukas Hoyer, Dengxin Dai, Luc Van GoolCVPR 2022 · 562 citations
Related papers
- UniDSeg: Unified Cross-Domain 3D Semantic Segmentation via Visual Foundation Models PriorYao Wu, Mingwei Xing, Yachao Zhang, Xiaotong Luo et al.NeurIPS 2024 · 15 citations
- Stronger, Fewer, & Superior: Harnessing Vision Foundation Models for Domain Generalized Semantic SegmentationZhixiang Wei, Lin Chen, Yi Jin, Xiaoxiao Ma et al.CVPR 2024 · 61 citations
- WildNet: Learning Domain Generalized Semantic Segmentation from the WildSuhyeon Lee, Hongje Seong, Seongwon Lee, Euntai KimCVPR 2022 · 95 citations
- Open-Vocabulary Domain Generalization in Urban-Scene SegmentationDong Zhao, Qi Zang, Nan Pu, Wenjing Li et al.CVPR 2026 · 3 citations
- Leveraging Depth and Language for Open-Vocabulary Domain-Generalized Semantic SegmentationSiyu Chen, Ting Han, Chengzheng Fu, Changshe Zhang et al.NeurIPS 2025 · 4 citations
