Your ViT is Secretly an Image Segmentation Model
Tommie Kerssies, Niccolò Cavagnero, Alexander Hermans, Narges Norouzi, Giuseppe Averta, Bastian Leibe, Gijs Dubbelman, Daan de Geus
摘要
Vision Transformers (ViTs) have shown remarkable performance and scalability across various computer vision tasks. To apply single-scale ViTs to image segmentation, existing methods adopt a convolutional adapter to generate multiscale features, a pixel decoder to fuse these features, and a Transformer decoder that uses the fused features to make predictions. In this paper, we show that the inductive biases introduced by these task-specific components can instead be learned by the ViT itself, given sufficiently large models and extensive pre-training. Based on these findings, we introduce the Encoder-only Mask Transformer (EoMT), which repurposes the plain ViT architecture to conduct image segmentation. With large-scale models and pre-training, EoMT obtains a segmentation accuracy similar to state-of-the-art models that use task-specific components. At the same time, EoMT is significantly faster than these methods due to its architectural simplicity, e.g., up to 4× faster with ViT-L. Across a range of model sizes, EoMT demonstrates an optimal balance between segmentation accuracy and prediction speed, suggesting that compute resources are better spent on scaling the ViT itself rather than adding architectural complexity. Code: https://www.tue-mps.org/eomt/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Revisiting 2D Foundation Models for Scalable 3D Medical Image ClassificationHan Liu, Bogdan Georgescu, Yanbo Zhang, Youngjin Yoo 等CVPR 2026 · 被引用 10 次
- VidEoMT: Your ViT is Secretly Also a Video Segmentation ModelNarges Norouzi, Idil Esen Zulfikar, Niccolò Cavagnero, Tommie Kerssies 等CVPR 2026 · 被引用 8 次
- Towards Implicit Aggregation: Robust Image Representation for Place Recognition in the Transformer EraFeng Lu, Tong Jin, Canming Ye, Xiangyuan Lan 等NeurIPS 2025 · 被引用 8 次
- Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and DatasetsMarianna Nezhurina, Tomer Porian, Giovanni Puccetti, Tommie Kerssies 等NeurIPS 2025 · 被引用 8 次
- Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT TransformersCris Claessens, Christiaan Viviers, Giacomo D'Amicantonio, Egor Bondarev 等CVPR 2026 · 被引用 6 次
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- Vision Transformer Adapter for Dense PredictionsZhe Chen, Yuchen Duan, Wenhai Wang, Junjun He 等ICLR 2023 · 被引用 204 次
- MPViT: Multi-Path Vision Transformer for Dense PredictionYoungwan Lee, Jonghee Kim, Jeffrey Willette, Sung Ju HwangCVPR 2022 · 被引用 339 次
- Turning Pre-Trained Vision Transformers into End-to-End Histopathology Whole Slide Image Models for Survival PredictionJiawen Li, Jiali Hu, Xitong Ling, Renao Yan 等CVPR 2026 · 被引用 1 次
- Auto-scaling Vision Transformers without TrainingWuyang Chen, Wei Huang, Xianzhi Du, Xiaodan Song 等ICLR 2022 · 被引用 27 次
- SegViT: Semantic Segmentation with Plain Vision TransformersBowen Zhang, Zhi Tian, Quan Tang, Xiangxiang Chu 等NeurIPS 2022 · 被引用 242 次
