v-CLR: View-Consistent Learning for Open-World Instance Segmentation
Chang-Bin Zhang, Jinhong Ni, Yujie Zhong, Kai Han
Abstract
In this paper, we address the challenging problem of openworld instance segmentation. Existing works have shown that vanilla visual networks are biased toward learning appearance information, e.g., texture, to recognize objects. This implicit bias causes the model to fail in detecting novel objects with unseen textures in the open-world setting. To address this challenge, we propose a learning framework, called view-Consistent LeaRning (v-CLR), which aims to enforce the model to learn appearance-invariant representations for robust instance segmentation. In v-CLR, we first introduce additional views for each image, where the texture undergoes significant alterations while preserving the image's underlying structure. We then encourage the model to learn the appearance-invariant representation by enforcing the consistency between object features across different views, for which we obtain class-agnostic object proposals using off-the-shelf unsupervised models that possess strong object-awareness. These proposals enable cross-view object feature matching, greatly reducing the appearance dependency while enhancing the objectawareness. We thoroughly evaluate our method on public benchmarks under both cross-class and cross-dataset settings, achieving state-of-the-art performance. Project
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Advancing Visual Large Language Model for Multi-Granular Versatile PerceptionWentao Xiang, Haoxian Tan, Yujie Zhong, Cong Wei et al.ICCV 2025 · 1 citation
- Prompt-Free Unknown Label Generation for Open World Detection in Remote SensingAbdullah Azeem, Ruisheng Wang, Qingquan Li, Abubakar SiddiqueCVPR 2026
- Transformers Trained via Gradient Descent Can Provably Learn a Class of Teacher ModelsChenyang Zhang, Qingyue Zhao, Quanquan Gu, Yuan CaoICLR 2026
Builds on36
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
Related papers
- Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask GuidancePhuc D. A. Nguyen, Tuan Duc Ngo, Evangelos Kalogerakis, Chuang Gan et al.CVPR 2024 · 45 citations
- Transformer-based Open-world Instance Segmentation with Cross-task Consistency RegularizationXizhe Xue, Dongdong Yu, Lingqiao Liu, Yu Liu et al.ACM MM 2023 · 2 citations
- Open-World Instance Segmentation: Exploiting Pseudo Ground Truth From Learned Pairwise AffinityWeiyao Wang, Matt Feiszli, Heng Wang, Jitendra Malik et al.CVPR 2022 · 39 citations
- SegPrompt: Boosting Open-world Segmentation via Category-level Prompt LearningMuzhi Zhu, Hengtao Li, Hao Chen, Chengxiang Fan et al.ICCV 2023 · 26 citations
- CrOC: Cross-View Online Clustering for Dense Visual Representation LearningThomas Stegmüller, Tim Lebailly, Behzad Bozorgtabar, Tinne Tuytelaars et al.CVPR 2023
