MobileInst: Video Instance Segmentation on the Mobile
Renhong Zhang, Tianheng Cheng, Shusheng Yang, Haoyi Jiang, Shuai Zhang, Jiancheng Lyu, Xin Li, Xiaowen Ying, Dashan Gao, Wenyu Liu, Xinggang Wang
摘要
Video instance segmentation on mobile devices is an important yet very challenging edge AI problem. It mainly suffers from (1) heavy computation and memory costs for frameby-frame pixel-level instance perception and (2) complicated heuristics for tracking objects. To address those issues, we present MobileInst, a lightweight and mobile-friendly framework for video instance segmentation on mobile devices. Firstly, MobileInst adopts a mobile vision transformer to extract multi-level semantic features and presents an efficient query-based dual-transformer instance decoder for mask kernels and a semantic-enhanced mask decoder to generate instance segmentation per frame. Secondly, MobileInst exploits simple yet effective kernel reuse and kernel association to track objects for video instance segmentation. Further, we propose temporal query passing to enhance the tracking ability for kernels. We conduct experiments on COCO and YouTube-VIS datasets to demonstrate the superiority of Mo-bileInst and evaluate the inference latency on one single CPU core of Snapdragon ® 778G Mobile Platform, without other methods of acceleration. On the COCO dataset, MobileInst achieves 31.2 mask AP and 433 ms on the mobile CPU, which reduces the latency by 50% compared to the previous SOTA. For video instance segmentation, MobileInst achieves 35.0 AP on YouTube-VIS 2019 and 30.1 AP on YouTube-VIS 2021. Code will be available to facilitate real-world applications and future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- RMP-SAM: Towards Real-Time Multi-Purpose Segment AnythingShilin Xu, Haobo Yuan, Qingyu Shi, Lu Qi 等ICLR 2025
- LightAVSeg: Lightweight Audio-Visual SegmentationQing Zhong, Guodong Ding, Lingqiao Liu, Zaiwen Feng 等ICML 2026
它引用的顶会 Paper25
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 被引用 2,162 次
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 被引用 2,075 次
- SOLOv2: Dynamic and Fast Instance SegmentationXinlong Wang, Rufeng Zhang, Tao Kong, Lei Li 等NeurIPS 2020 · 被引用 1,193 次
相关 Paper
- MinVIS: A Minimal Video Instance Segmentation Framework without Video-based TrainingDe-An Huang, Zhiding Yu, Anima AnandkumarNeurIPS 2022 · 被引用 135 次
- VITA: Video Instance Segmentation via Object Token AssociationMiran Heo, Sukjun Hwang, Seoung Wug Oh, Joon-Young Lee 等NeurIPS 2022 · 被引用 146 次
- InstanceFormer: An Online Video Instance Segmentation FrameworkRajat Koner, Tanveer Hannan, Suprosanna Shit, Sahand Sharifzadeh 等AAAI 2023 · 被引用 16 次
- Video Instance Segmentation using Inter-Frame Communication TransformersSukjun Hwang, Miran Heo, Seoung Wug Oh, Seon Joo KimNeurIPS 2021 · 被引用 174 次
- Instances as QueriesYuxin Fang, Shusheng Yang, Xinggang Wang, Yu Li 等ICCV 2021 · 被引用 331 次
