FastInst: A Simple Query-Based Model for Real-Time Instance Segmentation
Junjie He, Pengyu Li, Yifeng Geng, Xuansong Xie
Abstract
Recent attention in instance segmentation has focused on query-based models. Despite being non-maximum suppression (NMS)-free and end-to-end, the superiority of these models on high-accuracy real-time benchmarks has not been well demonstrated. In this paper, we show the strong potential of query-based models on efficient instance segmentation algorithm designs. We present FastInst, a simple, effective query-based framework for real-time instance segmentation. FastInst can execute at a real-time speed (i.e., 32.5 FPS) while yielding an AP of more than 40 (i.e., 40.5 AP) on COCO test-dev without bells and whistles. Specifically, FastInst follows the meta-architecture of recently introduced Mask2Former. Its key designs include instance activation-guided queries, dual-path update strategy, and ground truth mask-guided learning, which enable us to use lighter pixel decoders, fewer Transformer decoder layers, while achieving better performance. The experiments show that FastInst outperforms most state-of-the-art real-time counterparts, including strong fully convolutional baselines, in both speed and accuracy. Code can be found at https://github.com/junjiehe96/FastInst.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2fd69496-7db9-4739-99ee-ceef3e38b8c2Cited by top-tier papers13
- RF-DETR: Neural Architecture Search for Real-Time Detection TransformersIsaac Robinson, Peter Robicheaux, Matvei Popov, Deva Ramanan et al.ICLR 2026 · 161 citations
- MGQFormer: Mask-Guided Query-Based Transformer for Image Manipulation LocalizationKunlun Zeng, Ri Cheng, Weimin Tan, Bo YanAAAI 2024 · 23 citations
- DI-MaskDINO: A Joint Object Detection and Instance Segmentation ModelZhixiong Nan, Xianghong Li, Tao Xiang, Jifeng DaiNeurIPS 2024 · 15 citations
- ForestFormer3D: A Unified Framework for End-to-End Segmentation of Forest LiDAR 3D Point CloudsBinbin Xiang, Maciej Wielgosz, Stefano Puliti, Kamil Král et al.ICCV 2025 · 12 citations
- MobileInst: Video Instance Segmentation on the MobileRenhong Zhang, Tianheng Cheng, Shusheng Yang, Haoyi Jiang et al.AAAI 2024 · 10 citations
Builds on17
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 2,075 citations
- SOLOv2: Dynamic and Fast Instance SegmentationXinlong Wang, Rufeng Zhang, Tao Kong, Lei Li et al.NeurIPS 2020 · 1,193 citations
- TensorMask: A Foundation for Dense Object SegmentationXinlei Chen, Ross B. Girshick, Kaiming He, Piotr DollárICCV 2019 · 357 citations
Related papers
- Sparse Instance Activation for Real-Time Instance SegmentationTianheng Cheng, Xinggang Wang, Shaoyu Chen, Wenqiang Zhang et al.CVPR 2022 · 182 citations
- Instances as QueriesYuxin Fang, Shusheng Yang, Xinggang Wang, Yu Li et al.ICCV 2021 · 331 citations
- SOIT: Segmenting Objects with Instance-Aware TransformersXiaodong Yu, Dahu Shi, Xing Wei, Ye Ren et al.AAAI 2022 · 32 citations
- Masked-attention Mask Transformer for Universal Image SegmentationBowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov et al.CVPR 2022
- HyPiDecoder: Hybrid Pixel Decoder for Efficient Segmentation and DetectionFengzhe Zhou, Humphrey ShiICCV 2025 · 1 citation
