Learning Equivariant Segmentation with Instance-Unique Querying
Wenguan Wang, James Liang, Dongfang Liu
Abstract
Prevalent state-of-the-art instance segmentation methods fall into a query-based scheme, in which instance masks are derived by querying the image feature using a set of instance-aware embeddings. In this work, we devise a new training framework that boosts query-based models through discriminative query embedding learning. It explores two essential properties, namely dataset-level uniqueness and transformation equivariance, of the relation between queries and instances. First, our algorithm uses the queries to retrieve the corresponding instances from the whole training dataset, instead of only searching within individual scenes. As querying instances across scenes is more challenging, the segmenters are forced to learn more discriminative queries for effective instance separation. Second, our algorithm encourages both image (instance) representations and queries to be equivariant against geometric transformations, leading to more robust, instance-query matching. On top of four famous, query-based models ( CondInst, SOLOv2, SOTR, and Mask2Former), our training algorithm provides significant performance gains ( +1.6 - 3.2 AP) on COCO dataset. In addition, our algorithm promotes the performance of SOLOv2 by 2.7 AP, on LVISv1 dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 78143d0f-66d5-4ec6-8af0-949b415a86e8Cited by top-tier papers10
- E2VPT: An Effective and Efficient Approach for Visual Prompt TuningCheng Han, Qifan Wang, Yiming Cui, Zhiwen Cao et al.ICCV 2023 · 108 citations
- CLUSTSEG: Clustering for Universal SegmentationJames Chenhao Liang, Tianfei Zhou, Dongfang Liu, Wenguan WangICML 2023 · 85 citations
- ClusterFomer: Clustering As A Universal Visual LearnerJames Liang, Yiming Cui, Qifan Wang, Tong Geng et al.NeurIPS 2023 · 63 citations
- Large-Scale Person Detection and Localization using Overhead Fisheye CamerasLu Yang, Liulei Li, Xueshi Xin, Yifan Sun et al.ICCV 2023 · 35 citations
- Unified 3D Segmenter As Prototypical ClassifiersZheyun Qin, Cheng Han, Qifan Wang, Xiushan Nie et al.NeurIPS 2023 · 27 citations
Builds on27
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 2,075 citations
Related papers
- SOLQ: Segmenting Objects by Learning QueriesBin Dong, Fangao Zeng, Tiancai Wang, Xiangyu Zhang et al.NeurIPS 2021 · 143 citations
- SOIT: Segmenting Objects with Instance-Aware TransformersXiaodong Yu, Dahu Shi, Xing Wei, Ye Ren et al.AAAI 2022 · 32 citations
- FastInst: A Simple Query-Based Model for Real-Time Instance SegmentationJunjie He, Pengyu Li, Yifeng Geng, Xuansong XieCVPR 2023
- Instances as QueriesYuxin Fang, Shusheng Yang, Xinggang Wang, Yu Li et al.ICCV 2021 · 331 citations
- SOLOv2: Dynamic and Fast Instance SegmentationXinlong Wang, Rufeng Zhang, Tao Kong, Lei Li et al.NeurIPS 2020 · 1,193 citations
