Putting the Object Back into Video Object Segmentation
Ho Kei Cheng, Seoung Wug Oh, Brian L. Price, Joon-Young Lee, Alexander G. Schwing
Abstract
We present Cutie, a video object segmentation (VOS) network with object-level memory reading, which puts the object representation from memory back into the video object segmentation result. Recent works on VOS employ bottom-up pixel-level memory reading which struggles due to matching noise, especially in the presence of distractors, resulting in lower performance in more challenging data. In contrast, Cutie performs top-down object-level memory reading by adapting a small set of object queries. Via those, it interacts with the bottom-up pixel features iteratively with a query-based object transformer (qt, hence Cutie). The object queries act as a high-level summary of the target object, while high-resolution feature maps are retained for accurate segmentation. Together with foreground-background masked attention, Cutie cleanly separates the semantics of the foreground object from the background. On the challenging MOSE dataset, Cutie improves by 8.7 J &F over XMem with a similar running time and improves by 4.2 J &F over DeAOT while being three times faster. Code is available at: hkchengrex.github.io/Cutie.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 21780c28-fd5b-43f3-b862-65b4ad0d5aaeCited by top-tier papers61
- VR-GS: A Physical Dynamics-Aware Interactive Gaussian Splatting System in Virtual RealityYing Jiang, Chang Yu, Tianyi Xie, Xuan Li et al.SIGGRAPH 2024 · 153 citations
- DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous GraspingYifan Zhong, Xuchuan Huang, Ruochong Li, Ceyao Zhang et al.AAAI 2026 · 89 citations
- Segment Every Reference Object in Spatial and Temporal SpacesJiannan Wu, Yi Jiang, Bin Yan, Huchuan Lu et al.ICCV 2023 · 29 citations
- RMem: Restricted Memory Banks Improve Video Object SegmentationJunbao Zhou, Ziqi Pang, Yu-Xiong WangCVPR 2024 · 18 citations
- Advancing Complex Video Object Segmentation via Progressive Concept ConstructionZhixiong Zhang, Shuangrui Ding, Xiaoyi Dong, Songxin He et al.ICLR 2026 · 17 citations
Builds on44
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 403 citations
- Associating Objects with Transformers for Video Object SegmentationZongxin Yang, Yunchao Wei, Yi YangNeurIPS 2021 · 398 citations
Related papers
- Hierarchical Memory Matching Network for Video Object SegmentationHongje Seong, Seoung Wug Oh, Joon-Young Lee, Seongwon Lee et al.ICCV 2021 · 126 citations
- Look Before You Match: Instance Understanding Matters in Video Object SegmentationJunke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo et al.CVPR 2023
- Memory Aggregation Networks for Efficient Interactive Video Object SegmentationJiaxu Miao, Yunchao Wei, Yi YangCVPR 2020
- Object Guided External Memory Network for Video Object DetectionHanming Deng, Yang Hua, Tao Song, Zongpu Zhang et al.ICCV 2019 · 109 citations
- Shallow Features Matter: Hierarchical Memory with Heterogeneous Interaction for Unsupervised Video Object SegmentationXiangyu Zheng, Songcheng He, Wanyun Li, Xiaoqiang Li et al.ACM MM 2025
