PolyFormer: Referring Image Segmentation as Sequential Polygon Generation
Jiang Liu, Hui Ding, Zhaowei Cai, Yuting Zhang, Ravi Kumar Satzoda, Vijay Mahadevan, R. Manmatha
Abstract
In this work, instead of directly predicting the pixel-level segmentation masks, the problem of referring image segmentation is formulated as sequential polygon generation, and the predicted polygons can be later converted into segmentation masks. This is enabled by a new sequence-tosequence framework, Polygon Transformer (PolyFormer), which takes a sequence of image patches and text query tokens as input, and outputs a sequence of polygon vertices autoregressively. For more accurate geometric localization, we propose a regression-based decoder, which predicts the precise floating-point coordinates directly, without any coordinate quantization error. In the experiments, PolyFormer outperforms the prior art by a clear margin, e.g., 5.40% and 4.52% absolute improvements on the challenging Re-fCOCO+ and RefCOCOg datasets. It also shows strong generalization ability when evaluated on the referring video segmentation task without fine-tuning, e.g., achieving competitive 61.5% J &F on the Ref-DAVIS17 dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 361d1c33-1fb2-4bfe-9fd1-ec6c80700ce1Cited by top-tier papers75
- Segment Everything Everywhere All at OnceXueyan Zou, Jianwei Yang, Hao Zhang, Feng Li et al.NeurIPS 2023 · 889 citations
- SOC: Semantic-Assisted Object Cluster for Referring Video Object SegmentationZhuoyan Luo, Yicheng Xiao, Yong Liu, Shuyan Li et al.NeurIPS 2023 · 89 citations
- Hierarchical Open-vocabulary Universal Image SegmentationXudong Wang, Shufan Li, Konstantinos Kallidromitis, Yusuke Kato et al.NeurIPS 2023 · 74 citations
- Described Object Detection: Liberating Object Detection with Flexible ExpressionsChi Xie, Zhao Zhang, Yixuan Wu, Feng Zhu et al.NeurIPS 2023 · 69 citations
- SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal FusionMing Dai, Lingfeng Yang, Yihao Xu, Zhenhua Feng et al.NeurIPS 2024 · 67 citations
Builds on35
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 2,075 citations
Related papers
- Vision-Language Transformer and Query Generation for Referring SegmentationHenghui Ding, Chang Liu, Suchen Wang, Xudong JiangICCV 2021 · 359 citations
- Contrastive Grouping with Transformer for Referring Image SegmentationJiajin Tang, Ge Zheng, Cheng Shi, Sibei YangCVPR 2023
- Instance Segmentation with Mask-supervised Polygonal Boundary TransformersJustin Lazarow, Weijian Xu, Zhuowen TuCVPR 2022 · 55 citations
- Language as Queries for Referring Video Object SegmentationJiannan Wu, Yi Jiang, Peize Sun, Zehuan Yuan et al.CVPR 2022 · 143 citations
- ReSTR: Convolution-free Referring Image Segmentation Using TransformersNamyup Kim, Dongwon Kim, Suha Kwak, Cuiling Lan et al.CVPR 2022 · 149 citations
