PEM: Prototype-Based Efficient MaskFormer for Image Segmentation
Niccolò Cavagnero, Gabriele Rosi, Claudia Cuttano, Francesca Pistilli, Marco Ciccone, Giuseppe Averta, Fabio Cermelli
摘要
Recent transformer-based architectures have shown impressive results in the field of image segmentation. Thanks to their flexibility, they obtain outstanding performance in multiple segmentation tasks, such as semantic and panoptic, under a single unified framework. To achieve such impressive performance, these architectures employ intensive operations and require substantial computational resources, which are often not available, especially on edge devices. To fill this gap, we propose Prototype-based Efficient Mask-Former (PEM), an efficient transformer-based architecture that can operate in multiple segmentation tasks. PEM proposes a novel prototype-based cross-attention which leverages the redundancy of visual features to restrict the computation and improve the efficiency without harming the performance. In addition, PEM introduces an efficient multiscale feature pyramid network, capable of extracting features that have high semantic content in an efficient way, thanks to the combination of deformable convolutions and context-based self-modulation. We benchmark the proposed PEM architecture on two tasks, semantic and panoptic segmentation, evaluated on two different datasets, Cityscapes and ADE20K. PEM demonstrates outstanding performance on every task and dataset, outperforming task-specific architectures while being comparable and even better than computationally expensive baselines. Code is available at https://github.com/NiccoloCavagnero/PEM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- EOV-Seg: Efficient Open-Vocabulary Panoptic SegmentationHongwei Niu, Jie Hu, Jianghang Lin, Guannan Jiang 等AAAI 2025 · 被引用 11 次
- VidEoMT: Your ViT is Secretly Also a Video Segmentation ModelNarges Norouzi, Idil Esen Zulfikar, Niccolò Cavagnero, Tommie Kerssies 等CVPR 2026 · 被引用 8 次
- Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature AlignmentShi-Chen Zhang, Yunheng Li, Yu-Huan Wu, Qibin Hou 等ICCV 2025 · 被引用 8 次
- Refining Context-Entangled Content Segmentation via Curriculum Selection and Anti-Curriculum PromotionChunming He, Rihan Zhang, Fengyang Xiao, Dingming Zhang 等ICML 2026 · 被引用 7 次
- Feature Purification Matters: Suppressing Outlier Propagation for Training-Free Open-Vocabulary Semantic SegmentationShuo Jin, Siyue Yu, Bingfeng Zhang, Mingjie Sun 等ICCV 2025 · 被引用 3 次
它引用的顶会 Paper13
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 被引用 2,162 次
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 被引用 2,075 次
- Rethinking Vision Transformers for MobileNet Size and SpeedYanyu Li, Ju Hu, Yang Wen, Georgios Evangelidis 等ICCV 2023 · 被引用 300 次
- SwiftFormer: Efficient Additive Attention for Transformer-based Real-time Mobile Vision ApplicationsAbdelrahman M. Shaker, Muhammad Maaz, Hanoona Abdul Rasheed, Salman H. Khan 等ICCV 2023 · 被引用 213 次
相关 Paper
- Masked-attention Mask Transformer for Universal Image SegmentationBowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov 等CVPR 2022
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- SegNeXt: Rethinking Convolutional Attention Design for Semantic SegmentationMeng-Hao Guo, Cheng-Ze Lu, Qibin Hou, Zhengning Liu 等NeurIPS 2022 · 被引用 1,385 次
- RTFormer: Efficient Design for Real-Time Semantic Segmentation with TransformerJian Wang, Chenhui Gou, Qiman Wu, Haocheng Feng 等NeurIPS 2022 · 被引用 207 次
- SOTR: Segmenting Objects with TransformersRuohao Guo, Dantong Niu, Liao Qu, Zhenbo LiICCV 2021 · 被引用 123 次
