Learning to Segment Every Referring Object Point by Point
Mengxue Qu, Yu Wu, Yunchao Wei, Wu Liu, Xiaodan Liang, Yao Zhao
Abstract
Referring Expression Segmentation (RES) can facilitate pixel-level semantic alignment between vision and language. Most of the existing RES approaches require massive pixel-level annotations, which are expensive and exhaustive. In this paper, we propose a new partially supervised training paradigm for RES, i.e., training using abundant referring bounding boxes and only a few (e.g., 1%) pixel-level referring masks. To maximize the transferability from the REC model, we construct our model based on the point-based sequence prediction model. We propose the co-content teacher-forcing to make the model explicitly associate the point coordinates (scale values) with the referred spatial features, which alleviates the exposure bias caused by the limited segmentation masks. To make the most of referring bounding box annotations, we further propose the resampling pseudo points strategy to select more accurate pseudo-points as supervision. Extensive experiments show that our model achieves 52.06% in terms of accuracy (versus 58.93% in fully supervised setting) on Re-fCOCO+@testA, when only using 1% of the mask annotations. Code is available at https://github.com/ qumengxue/Partial-RES.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 71286495-e202-4d1c-a434-31e43d98a809Cited by top-tier papers2
- Decoupling Static and Hierarchical Motion Perception for Referring Video SegmentationShuting He, Henghui DingCVPR 2024
- Open-Vocabulary Segmentation with Semantic-Assisted CalibrationYong Liu, Sule Bai, Guanbin Li, Yitong Wang et al.CVPR 2024
Builds on25
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang et al.ICCV 2019 · 441 citations
- Pix2seq: A Language Modeling Framework for Object DetectionTing Chen, Saurabh Saxena, Lala Li, David J. Fleet et al.ICLR 2022 · 435 citations
- Fast Convergence of DETR with Spatially Modulated Co-AttentionPeng Gao, Minghang Zheng, Xiaogang Wang, Jifeng Dai et al.ICCV 2021 · 392 citations
Related papers
- RefTeacher: A Strong Baseline for Semi-Supervised Referring Expression ComprehensionJiamu Sun, Gen Luo, Yiyi Zhou, Xiaoshuai Sun et al.CVPR 2023
- Unveiling Parts Beyond Objects: Towards Finer-Granularity Referring Expression SegmentationWenxuan Wang, Tongtian Yue, Yisi Zhang, Longteng Guo et al.CVPR 2024 · 5 citations
- AMLRIS: Alignment-aware Masked Learning for Referring Image SegmentationTongfei Chen, Shuo Yang, Yuguang Yang, Linlin Yang et al.ICLR 2026 · 1 citation
- Advancing Referring Expression Segmentation Beyond Single ImageYixuan Wu, Zhao Zhang, Chi Xie, Feng Zhu et al.ICCV 2023 · 25 citations
- Referring Image Segmentation Using Text SupervisionFang Liu, Yuhao Liu, Yuqiu Kong, Ke Xu et al.ICCV 2023 · 52 citations
