Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation
Manuel Traub, Martin V Butz
摘要
Current state-of-the-art segmentation models encode entire images before focusing on specific objects. This wastes computational resources. We introduce FLIP (Fovea-Like Input Patching), a parameter-efficient vision model that realizes object segmentation through biologically-inspired top-down attention. FLIP selectively samples multi-resolution patches centered on objects of interest from the input. As a result, it allocates high-resolution processing to object centers while maintaining coarser peripheral context. This off-grid, scale-invariant design enables FLIP to outperform META's Segment Anything models (SAM, SAM2 and fast variants) by large margins: With more than 440 fewer parameters, FLIP-Tiny (0.51M parameters) reaches a mean IoU of 79.90% while SAM2-L reaches 75.87% IoU (224.45M parameters). FLIP-Large even achieves 83.26% mean IoU (96.6M parameters), still running about faster than SAM2-L. We evaluate on six benchmarks in total. In five established benchmarks (Hypersim, KITTI-360, OpenImages, COCO, LVIS) FLIP consistently outperforms SAM and various variants of it. In our novel ObjaScale dataset, which stress-tests scale invariance with objects ranging from 0.0001% up to 25% of the image area, we show that FLIP segments even very small objects accurately, where existing models fail severely. FLIP opens new possibilities for real-time, object-centric vision and offers much higher energy efficiency. We believe that FLIP can act as a powerful foundation model, as it is very well-suited to track objects over time, for example, when being integrated into slot-based scene segmentation architectures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran 等NeurIPS 2020 · 被引用 1,275 次
相关 Paper
- Segment This Thing: Foveated Tokenization for Efficient Point-Prompted SegmentationTanner Schmidt, Richard A. NewcombeCVPR 2025
- Efficient-SAM2: Accelerating SAM2 with Object-Aware Visual Encoding and Memory RetrievalJing Zhang, Zhikai Li, Xuewen Liu, Qingyi GuICLR 2026 · 被引用 5 次
- Vision Transformer with Sparse Scan PriorYuguang Zhang, Qihang Fan, Huaibo HuangACM MM 2025 · 被引用 2 次
- Parameter-Inverted Image Pyramid NetworksXizhou Zhu, Xue Yang, Zhaokai Wang, Hao Li 等NeurIPS 2024 · 被引用 11 次
- AutoFocus: Efficient Multi-Scale InferenceMahyar Najibi, Bharat Singh, Larry DavisICCV 2019 · 被引用 143 次
