Better to Teach than to Give: Domain Generalized Semantic Segmentation via Agent Queries with Diffusion Model Guidance
Fan Li, Xuan Wang, Min Qi, Zhaoxiang Zhang, Yuelei Xu
Abstract
Domain Generalized Semantic Segmentation (DGSS) trains a model on a labeled source domain to generalize to unseen target domains with consistent contextual distribution and varying visual appearance. Most existing methods rely on domain randomization or data generation but struggle to capture the underlying scene distribution, resulting in the loss of useful semantic information. Inspired by the diffusion model's capability to generate diverse variations within a given scene context, we consider harnessing its rich prior knowledge of scene distribution to tackle the challenging DGSS task. In this paper, we propose a novel agent Query-driven learning framework based on Diffusion model guidance for DGSS, named QueryDiff. Our recipe comprises three key ingredients: (1) generating agent queries from segmentation features to aggregate semantic information about instances within the scene; (2) learning the inherent semantic distribution of the scene through agent queries guided by diffusion features; (3) refining segmentation features using optimized agent queries for robust mask predictions. Extensive experiments across various settings demonstrate that our method significantly outperforms previous state-of-the-art methods. Notably, it enhances the model's ability to generalize effectively to extreme domains, such as cubist art styles. Code is available at https: //github.com/FanLiHub/QueryDiff .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e721fb7-1c7d-494a-8d66-bf05e2a05a03Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- No Object Is an Island: Enhancing 3D Semantic Segmentation Generalization with Diffusion ModelsFan Li, Xuan Wang, Xuanbin Wang, Zhaoxiang Zhang et al.NeurIPS 2025 · 4 citations
- Collaborating Foundation Models for Domain Generalized Semantic SegmentationYasser Benigmim, Subhankar Roy, Slim Essid, Vicky Kalogeiton et al.CVPR 2024
- Exploring Semantic Consistency and Style Diversity for Domain Generalized Semantic SegmentationHongwei Niu, Linhuang Xie, Jianghang Lin, Shengchuan ZhangAAAI 2025 · 16 citations
- Exploring Probabilistic Modeling Beyond Domain Generalization for Semantic SegmentationI-Hsiang Chen, Hua-En Chang, Wei-Ting Chen, Jenq-Neng Hwang et al.ICCV 2025 · 2 citations
- Images as Noisy Labels: Unleashing the Potential of the Diffusion Model for Open-Vocabulary Semantic SegmentationFan Li, Xuanbin Wang, Xuan Wang, Zhaoxiang Zhang et al.ICCV 2025 · 3 citations
