DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception
Junjie Wang, Bin Chen, Yulin Li, Bin Kang, Yichi Chen, Zhuotao Tian
摘要
Dense visual prediction tasks have been constrained by their reliance on predefined categories, limiting their applicability in real-world scenarios where visual concepts are unbounded. While Vision-Language Models (VLMs) like CLIP have shown promise in open-vocabulary tasks, their direct application to dense prediction often leads to suboptimal performance due to limitations in local feature representation. In this work, we present our observation that CLIP's image tokens struggle to effectively aggregate information from spatially or semantically related regions, resulting in features that lack local discriminability and spatial consistency. To address this issue, we propose DeCLIP, a novel framework that enhances CLIP by decoupling the self-attention module to obtain "content" and "context" features respectively. The "content" features are aligned with image crop representations to improve local discriminability, while "context" features learn to retain the spatial correlations under the guidance of vision foundation models, such as DINO. Extensive experiments demonstrate that DeCLIP significantly outperforms existing methods across multiple open-vocabulary dense prediction tasks, including object detection and semantic segmentation. Code is available at https://github.com/xiaomoguhz/DeCLIP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe PriorYulin Li, Haokun Gui, Ziyang Fan, Junjie Wang 等NeurIPS 2025 · 被引用 18 次
- Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMsShangpin Peng, Weinong Wang, Zhuotao Tian, Senqiao Yang 等ICLR 2026 · 被引用 10 次
- Efficient Reasoning with Balanced ThinkingYulin Li, Tengyao Tu, Li Ding, Junjie Wang 等ICLR 2026 · 被引用 7 次
- Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLMTianyi Wu, Xiaoxi Sun, Yanhua Jiao, Yulin Li 等ICLR 2026 · 被引用 6 次
- PCA-Seg: Revisiting Cost Aggregation for Open-Vocabulary Semantic and Part SegmentationJianjian Yin, Tao Chen, Yi Chen, Gensheng Pei 等CVPR 2026 · 被引用 6 次
它引用的顶会 Paper40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
相关 Paper
- Unbiased Region-Language Alignment for Open-Vocabulary Dense PredictionYunheng Li, Yuxuan Li, Quan-Sheng Zeng, Wenhai Wang 等ICCV 2025 · 被引用 3 次
- ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense PredictionJuan Yeo, Soonwoo Cha, Jiwoo Song, Hyunbin Jin 等ICCV 2025
- CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense PredictionSize Wu, Wenwei Zhang, Lumin Xu, Sheng Jin 等ICLR 2024 · 被引用 129 次
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware PromptingYongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang 等CVPR 2022 · 被引用 527 次
- ResCLIP: Residual Attention for Training-free Dense Vision-language InferenceYuhang Yang, Jinhong Deng, Wen Li, Lixin DuanCVPR 2025
