CMT-DeepLab: Clustering Mask Transformers for Panoptic Segmentation
Qihang Yu, Huiyu Wang, Dahun Kim, Siyuan Qiao, Maxwell D. Collins, Yukun Zhu, Hartwig Adam, Alan L. Yuille, Liang-Chieh Chen
Abstract
We propose Clustering Mask Transformer (CMT-DeepLab), a transformer-based framework for panoptic segmentation designed around clustering. It rethinks the existing transformer architectures used in segmentation and detection; CMT-DeepLab considers the object queries as cluster centers, which fill the role of grouping the pixels when applied to segmentation. The clustering is computed with an alternating procedure, by first assigning pixels to the clusters by their feature affinity, and then updating the cluster centers and pixel features. Together, these operations comprise the Clustering Mask Transformer (CMT) layer, which produces cross-attention that is denser and more consistent with the final segmentation task. CMT-DeepLab improves the performance over prior art significantly by 4.4% PQ, achieving a new state-of-the-art of 55.7% PQ on the COCO test-dev set.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fcac905e-ec64-46f6-8d07-be499d30b65dCited by top-tier papers27
- An Image is Worth 32 Tokens for Reconstruction and GenerationQihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen et al.NeurIPS 2024 · 331 citations
- Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIPQihang Yu, Ju He, Xueqing Deng, Xiaohui Shen et al.NeurIPS 2023 · 285 citations
- Image as Set of PointsXu Ma, Yuqian Zhou, Huan Wang, Can Qin et al.ICLR 2023 · 221 citations
- OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and UnderstandingTao Zhang, Xiangtai Li, Hao Fei, Haobo Yuan et al.NeurIPS 2024 · 186 citations
- A Generalist Framework for Panoptic Segmentation of Images and VideosTing Chen, Lala Li, Saurabh Saxena, Geoffrey E. Hinton et al.ICCV 2023 · 140 citations
Builds on24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
Related papers
- CLUSTSEG: Clustering for Universal SegmentationJames Chenhao Liang, Tianfei Zhou, Dongfang Liu, Wenguan WangICML 2023 · 85 citations
- MaX-DeepLab: End-to-End Panoptic Segmentation With Mask TransformersHuiyu Wang, Yukun Zhu, Hartwig Adam, Alan L. Yuille et al.CVPR 2021
- Masked-attention Mask Transformer for Universal Image SegmentationBowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov et al.CVPR 2022
- Panoptic SegFormer: Delving Deeper into Panoptic Segmentation with TransformersZhiqi Li, Wenhai Wang, Enze Xie, Zhiding Yu et al.CVPR 2022 · 145 citations
- ClusterFomer: Clustering As A Universal Visual LearnerJames Liang, Yiming Cui, Qifan Wang, Tong Geng et al.NeurIPS 2023 · 63 citations
