DPSeg: Dual-Prompt Cost Volume Learning for Open-Vocabulary Semantic Segmentation
Ziyu Zhao, Xiaoguang Li, Lingjia Shi, Nasrin Imanpour, Song Wang
摘要
Open-vocabulary semantic segmentation aims to segment images into distinct semantic regions for both seen and unseen categories at the pixel level. Current methods utilize text embeddings from pre-trained vision-language models like CLIP but struggle with the inherent domain gap between image and text embeddings, even after extensive alignment during training. Additionally, relying solely on deep text-aligned features limits shallow-level feature guidance, which is crucial for detecting small objects and fine details, ultimately reducing segmentation accuracy. To address these limitations, we propose a dual prompting framework, DPSeg, for this task. Our approach combines dualprompt cost volume generation, a cost volume-guided decoder, and a semantic-guided prompt refinement strategy that leverages our dual prompting scheme to mitigate alignment issues in visual prompt generation. By incorporating visual embeddings from a visual prompt encoder, our approach reduces the domain gap between text and image embeddings while providing multi-level guidance through shallow features. Extensive experiments demonstrate that our method significantly outperforms existing state-of-theart approaches on multiple public datasets. The code and models are available: https://github.com/ ziyuz-vision/DPSeg .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- PANet: Few-Shot Image Semantic Segmentation With Prototype AlignmentKaixin Wang, Jun Hao Liew, Yingtian Zou, Daquan Zhou 等ICCV 2019 · 被引用 1,404 次
相关 Paper
- Dual Semantic Guidance for Open Vocabulary Semantic SegmentationZhengyang Wang, Tingliang Feng, Fan Lyu, Fanhua Shang 等CVPR 2025
- Open-Vocabulary Semantic Segmentation with Mask-adapted CLIPFeng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li 等CVPR 2023
- PhaseAlign: Complex Phase Alignment for Stable Open-Vocabulary Semantic SegmentationJiankang Wang, Dingding Jia, Zhoushuopeng, Xuan WangICML 2026
- LoGoSeg: Integrating Local and Global Features for Open-Vocabulary Semantic SegmentationJunyang Chen, Xiangbo Lv, Zhiqiang Kou, Xingdong Sheng 等AAAI 2026
- CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic SegmentationSeokju Cho, Heeseong Shin, Sunghwan Hong, Anurag Arnab 等CVPR 2024
