Putting the Object Back into Video Object Segmentation
Ho Kei Cheng, Seoung Wug Oh, Brian L. Price, Joon-Young Lee, Alexander G. Schwing
摘要
We present Cutie, a video object segmentation (VOS) network with object-level memory reading, which puts the object representation from memory back into the video object segmentation result. Recent works on VOS employ bottom-up pixel-level memory reading which struggles due to matching noise, especially in the presence of distractors, resulting in lower performance in more challenging data. In contrast, Cutie performs top-down object-level memory reading by adapting a small set of object queries. Via those, it interacts with the bottom-up pixel features iteratively with a query-based object transformer (qt, hence Cutie). The object queries act as a high-level summary of the target object, while high-resolution feature maps are retained for accurate segmentation. Together with foreground-background masked attention, Cutie cleanly separates the semantics of the foreground object from the background. On the challenging MOSE dataset, Cutie improves by 8.7 J &F over XMem with a similar running time and improves by 4.2 J &F over DeAOT while being three times faster. Code is available at: hkchengrex.github.io/Cutie.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper61
- VR-GS: A Physical Dynamics-Aware Interactive Gaussian Splatting System in Virtual RealityYing Jiang, Chang Yu, Tianyi Xie, Xuan Li 等SIGGRAPH 2024 · 被引用 153 次
- DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous GraspingYifan Zhong, Xuchuan Huang, Ruochong Li, Ceyao Zhang 等AAAI 2026 · 被引用 89 次
- Segment Every Reference Object in Spatial and Temporal SpacesJiannan Wu, Yi Jiang, Bin Yan, Huchuan Lu 等ICCV 2023 · 被引用 29 次
- RMem: Restricted Memory Banks Improve Video Object SegmentationJunbao Zhou, Ziqi Pang, Yu-Xiong WangCVPR 2024 · 被引用 18 次
- Advancing Complex Video Object Segmentation via Progressive Concept ConstructionZhixiong Zhang, Shuangrui Ding, Xiaoyi Dong, Songxin He 等ICLR 2026 · 被引用 17 次
它引用的顶会 Paper44
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 被引用 845 次
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 被引用 403 次
- Associating Objects with Transformers for Video Object SegmentationZongxin Yang, Yunchao Wei, Yi YangNeurIPS 2021 · 被引用 398 次
相关 Paper
- Hierarchical Memory Matching Network for Video Object SegmentationHongje Seong, Seoung Wug Oh, Joon-Young Lee, Seongwon Lee 等ICCV 2021 · 被引用 126 次
- Look Before You Match: Instance Understanding Matters in Video Object SegmentationJunke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo 等CVPR 2023
- Memory Aggregation Networks for Efficient Interactive Video Object SegmentationJiaxu Miao, Yunchao Wei, Yi YangCVPR 2020
- Object Guided External Memory Network for Video Object DetectionHanming Deng, Yang Hua, Tao Song, Zongpu Zhang 等ICCV 2019 · 被引用 109 次
- Shallow Features Matter: Hierarchical Memory with Heterogeneous Interaction for Unsupervised Video Object SegmentationXiangyu Zheng, Songcheng He, Wanyun Li, Xiaoqiang Li 等ACM MM 2025
