Learning Frequency-Adapted Vision Foundation Model for Domain Generalized Semantic Segmentation
Qi Bi, Jingjun Yi, Hao Zheng, Haolan Zhan, Yawen Huang, Wei Ji, Yuexiang Li, Yefeng Zheng
摘要
The emerging vision foundation model (VFM) has inherited the ability to generalize to unseen images. Nevertheless, the key challenge of domain-generalized semantic segmentation (DGSS) lies in the domain gap attributed to the cross-domain styles, e.g., the variance of urban landscape and environment dependencies. Hence, maintaining the style-invariant property with varying domain styles becomes the key bottleneck in harnessing VFM for DGSS. The frequency space after Haar wavelet transform provides a feasible way to decouple the style information from the domain-invariant content, since the content and style information is retained in the low-and high-frequency components of the space, respectively. To this end, we propose a novel Frequency-Adapted (FADA) learning scheme to advance the frontier. Its overall idea is to separately tackle the content and style information by frequency tokens throughout the learning process. Particularly, the proposed FADA consists of two branches, i.e., low-and high-frequency branches. The former is able to stabilize the scene content, while the latter learns the scene styles and eliminates its impact to DGSS. Experiments conducted on various DGSS settings show the state-of-the-art performance of our FADA and its versatility to a variety of VFMs. Source code is available at https://github.com/BiQiWHU/FADA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Learning Fine-grained Domain Generalization via Hyperbolic State Space HallucinationQi Bi, Jingjun Yi, Haolan Zhan, Wei Ji 等AAAI 2025 · 被引用 8 次
- Leveraging Depth and Language for Open-Vocabulary Domain-Generalized Semantic SegmentationSiyu Chen, Ting Han, Chengzheng Fu, Changshe Zhang 等NeurIPS 2025 · 被引用 4 次
- Stronger, Steadier & Superior: Geometric Consistency in Depth VFM Forges Domain Generalized Semantic SegmentationSiyu Chen, Ting Han, Changshe Zhang, Xin Luo 等ICCV 2025 · 被引用 4 次
- A Simple Yet Mighty Hartley Diffusion Versatilist for Generalizable Dense Vision TasksQi Bi, Jingjun Yi, Huimin Huang, Hao Zheng 等ICCV 2025 · 被引用 3 次
- Local Precise Refinement: A Dual-Gated Mixture-of-Experts for Enhancing Foundation Model Generalization against Spectral ShiftsXi Chen, Maojun Zhang, Yu Liu, Shen YanCVPR 2026 · 被引用 3 次
它引用的顶会 Paper52
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang 等NeurIPS 2022 · 被引用 1,291 次
- ACDC: The Adverse Conditions Dataset with Correspondences for Semantic Driving Scene UnderstandingChristos Sakaridis, Dengxin Dai, Luc Van GoolICCV 2021 · 被引用 655 次
相关 Paper
- Learning Spectral-Decomposited Tokens for Domain Generalized Semantic SegmentationJingjun Yi, Qi Bi, Hao Zheng, Haolan Zhan 等ACM MM 2024 · 被引用 25 次
- AdaDCP: Learning an Adapter with Discrete Cosine Prior for Clear-to-Adverse Domain GeneralizationQi Bi, Yixian Shen, Jingjun Yi, Gui-Song XiaICCV 2025 · 被引用 7 次
- Causal-Tune: Mining Causal Factors from Vision Foundation Models for Domain Generalized Semantic SegmentationYin Zhang, Yongqiang Zhang, Yaoyue Zheng, Bogdan Raducanu 等AAAI 2026
- Learning Generalized Segmentation for Foggy-Scenes by Bi-directional Wavelet GuidanceQi Bi, Shaodi You, Theo GeversAAAI 2024 · 被引用 45 次
- FSDR: Frequency Space Domain Randomization for Domain GeneralizationJiaxing Huang, Dayan Guan, Aoran Xiao, Shijian LuCVPR 2021
