Learning Content-Enhanced Mask Transformer for Domain Generalized Urban-Scene Segmentation
Qi Bi, Shaodi You, Theo Gevers
摘要
Domain-generalized urban-scene semantic segmentation (USSS) aims to learn generalized semantic predictions across diverse urban-scene styles. Unlike generic domain gap challenges, USSS is unique in that the semantic categories are often similar in different urban scenes, while the styles can vary significantly due to changes in urban landscapes, weather conditions, lighting, and other factors. Existing approaches typically rely on convolutional neural networks (CNNs) to learn the content of urban scenes.
In this paper, we propose a Content-enhanced Mask TransFormer (CMFormer) for domain-generalized USSS. The main idea is to enhance the focus of the fundamental component, the mask attention mechanism, in Transformer segmentation models on content information. We have observed through empirical analysis that a mask representation effectively captures pixel segments, albeit with reduced robustness to style variations. Conversely, its lower-resolution counterpart exhibits greater ability to accommodate style variations, while being less proficient in representing pixel segments. To harness the synergistic attributes of these two approaches, we introduce a novel content-enhanced mask attention mechanism. It learns mask queries from both the image feature and its down-sampled counterpart, aiming to simultaneously encapsulate the content and address stylistic variations. These features are fused into a Transformer decoder and integrated into a multi-resolution content-enhanced mask attention learning scheme.
Extensive experiments conducted on various domain-generalized urban-scene segmentation datasets demonstrate that the proposed CMFormer significantly outperforms existing CNN-based methods by up to 14.0% mIoU and the contemporary HGFormer by up to 1.7% mIoU. The source code is publicly available at https://github.com/BiQiWHU/CMFormer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Learning Frequency-Adapted Vision Foundation Model for Domain Generalized Semantic SegmentationQi Bi, Jingjun Yi, Hao Zheng, Haolan Zhan 等NeurIPS 2024 · 被引用 62 次
- Learning Generalized Segmentation for Foggy-Scenes by Bi-directional Wavelet GuidanceQi Bi, Shaodi You, Theo GeversAAAI 2024 · 被引用 45 次
- Learning Generalized Medical Image Segmentation from Decoupled Feature QueriesQi Bi, Jingjun Yi, Hao Zheng, Wei Ji 等AAAI 2024 · 被引用 41 次
- Learning Spectral-Decomposited Tokens for Domain Generalized Semantic SegmentationJingjun Yi, Qi Bi, Hao Zheng, Haolan Zhan 等ACM MM 2024 · 被引用 25 次
- Exploring Semantic Consistency and Style Diversity for Domain Generalized Semantic SegmentationHongwei Niu, Linhuang Xie, Jianghang Lin, Shengchuan ZhangAAAI 2025 · 被引用 16 次
它引用的顶会 Paper27
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- Segmenter: Transformer for Semantic SegmentationRobin Strudel, Ricardo Garcia, Ivan Laptev, Cordelia SchmidICCV 2021 · 被引用 1,898 次
- ACDC: The Adverse Conditions Dataset with Correspondences for Semantic Driving Scene UnderstandingChristos Sakaridis, Dengxin Dai, Luc Van GoolICCV 2021 · 被引用 655 次
相关 Paper
- HGFormer: Hierarchical Grouping Transformer for Domain Generalized Semantic SegmentationJian Ding, Nan Xue, Gui-Song Xia, Bernt Schiele 等CVPR 2023
- WildNet: Learning Domain Generalized Semantic Segmentation from the WildSuhyeon Lee, Hongje Seong, Seongwon Lee, Euntai KimCVPR 2022 · 被引用 95 次
- Exploiting Domain Properties in Language-Driven Domain Generalization for Semantic SegmentationSeogkyu Jeon, Kibeom Hong, Hyeran ByunICCV 2025 · 被引用 2 次
- Scaling up Image Segmentation across Data and TasksPei Wang, Zhaowei Cai, Hao Yang, Ashwin Swaminathan 等CVPR 2025
- DAFormer: Improving Network Architectures and Training Strategies for Domain-Adaptive Semantic SegmentationLukas Hoyer, Dengxin Dai, Luc Van GoolCVPR 2022 · 被引用 562 次
