Lune

ICCV2025Top-tier venue

Unbiased Region-Language Alignment for Open-Vocabulary Dense Prediction

Yunheng Li, Yuxuan Li, Quan-Sheng Zeng, Wenhai Wang, Qibin Hou, Ming-Ming Cheng

2025Year
3Citations
10Top-tier citations

Abstract

Pre-trained vision-language models (VLMs), such as CLIP, have demonstrated impressive zero-shot recognition capability, but still underperform in dense prediction tasks. Selfdistillation recently is emerging as a promising approach for fine-tuning VLMs to better adapt to local regions without requiring extensive annotations. However, previous stateof-the-art approaches often suffer from significant 'foreground bias', where models tend to wrongly identify background regions as foreground objects. To alleviate this issue, we propose DenseVLM, a framework designed to learn unbiased region-language alignment from powerful pretrained VLM representations. DenseVLM leverages the pretrained VLM to retrieve categories for unlabeled regions and then decouples the interference between foreground and background features. This separation ensures accurate region-category alignment while maintaining semantic distinctions during training. We show that DenseVLM can directly replace the original VLM in open-vocabulary object detection and image segmentation methods, leading to notable performance improvements. Furthermore, it exhibits promising zero-shot scalability when training on more extensive and diverse datasets. Our code is publicly available https://github.com/HVision-NKU/ DenseVLM .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext b9a9bd57-5b84-4ded-840f-c7d0a6d05260

Cited by top-tier papers10

Ask how each one uses it

Builds on36

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines