PhaseAlign: Complex Phase Alignment for Stable Open-Vocabulary Semantic Segmentation
Jiankang Wang, Dingding Jia, Zhoushuopeng, Xuan Wang
Abstract
Open-Vocabulary Segmentation(OVS) aims to achieve pixel-level semantic recognition from arbitrary text queries. Existing large-scale visual-linguistic models, such as CLIP, perform well in zero-shot generalization, but their image-level training objectives and real-valued cross-modal alignment mix amplitude and phase information, limiting fine-grained segmentation and often causing blurred boundaries and fragmented structures. Inspired by the ability of electromagnetic wave phase to control interference independently of amplitude, we propose PhaseAlign, an OVS framework based on Complex Phase Alignment (CPA). CPA explicitly decouples the magnitude and phase of visual and textual embeddings in the complex domain, refining effective features for stable cross-modal alignment. To further enhance structural awareness, we introduce spatial-aware cross-modal projection, which models local neighborhood relations via multi-scale spatial contrast normalization, and attention-guided affinity modeling, which leverages pre-trained ViT self-attention to propagate category activations, improving boundary clarity and region integrity. Experiments show that PhaseAlign achieves state-of-the-art performance on multiple OVS benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 535cb896-c9bb-4084-b107-e5fc5a8593bfBuilds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
- GroupViT: Semantic Segmentation Emerges from Text SupervisionJiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon et al.CVPR 2022 · 398 citations
Related papers
- DPSeg: Dual-Prompt Cost Volume Learning for Open-Vocabulary Semantic SegmentationZiyu Zhao, Xiaoguang Li, Lingjia Shi, Nasrin Imanpour et al.CVPR 2025
- InfoCLIP: Bridging Vision-Language Pretraining and Open-Vocabulary Semantic Segmentation via Information-Theoretic Alignment TransferMuyao Yuan, Yuanhong Zhang, Weizhan Zhang, Lan Ma et al.AAAI 2026 · 1 citation
- Dual Semantic Guidance for Open Vocabulary Semantic SegmentationZhengyang Wang, Tingliang Feng, Fan Lyu, Fanhua Shang et al.CVPR 2025
- LoGoSeg: Integrating Local and Global Features for Open-Vocabulary Semantic SegmentationJunyang Chen, Xiangbo Lv, Zhiqiang Kou, Xingdong Sheng et al.AAAI 2026
- CP-CLIP: Customized Parameter Generation for Open-vocabulary Semantic SegmentationZelin Peng, Zhengqin Xu, Feilong Tang, Wei ShenAAAI 2026
