Video State-Changing Object Segmentation
Jiangwei Yu, Xiang Li, Xinran Zhao, Hongming Zhang, Yu-Xiong Wang
摘要
Daily objects commonly experience state changes. For example, slicing a cucumber changes its state from whole to sliced. Learning about object state changes in Video Object Segmentation (VOS) is crucial for understanding and interacting with the visual world. Conventional VOS benchmarks do not consider this challenging yet crucial problem. This paper makes a pioneering effort to introduce a weakly-supervised benchmark on Video State-Changing Object Segmentation (VSCOS). We construct our VSCOS benchmark by selecting state-changing videos from existing datasets. In advocate of an annotation-efficient approach towards state-changing object segmentation, we only annotate the first and last frames of training videos, which is different from conventional VOS. Notably, an open-vocabulary setting is included to evaluate the generalization to novel types of objects or state changes. We empirically illustrate that state-of-the-art VOS models struggle with state-changing objects and lose track after the state changes. We analyze the main difficulties of our VSCOS task and identify three technical improvements, namely, fine-tuning strategies, representation learning, and integrating motion information. Applying these improvements results in a strong baseline for segmenting state-changing objects consistently. Our benchmark and baseline methods are publicly available at https://github.com/venom12138/VSCOS .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- RMem: Restricted Memory Banks Improve Video Object SegmentationJunbao Zhou, Ziqi Pang, Yu-Xiong WangCVPR 2024 · 被引用 18 次
- Tracking and Understanding Object TransformationsYihong Sun, Xinyu Yang, Jennifer J. Sun, Bharath HariharanNeurIPS 2025 · 被引用 5 次
- Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal InterventionHaijing Liu, Zhiyuan Song, Hefeng Wu, Tao Pu 等NeurIPS 2025 · 被引用 2 次
- Live Interactive Training for Video SegmentationXinyu Yang, Haozheng Yu, Yihong Sun, Bharath Hariharan 等CVPR 2026 · 被引用 1 次
- M^3-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object SegmentationZixuan Chen, Jiaxin Li, Junxuan Liang, Liming Tan 等CVPR 2025
它引用的顶会 Paper7
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 被引用 845 次
- Associating Objects with Transformers for Video Object SegmentationZongxin Yang, Yunchao Wei, Yi YangNeurIPS 2021 · 被引用 398 次
- Space-Time Correspondence as a Contrastive Random WalkAllan Jabri, Andrew Owens, Alexei A. EfrosNeurIPS 2020 · 被引用 356 次
- Decoupling Features in Hierarchical Propagation for Video Object SegmentationZongxin Yang, Yi YangNeurIPS 2022 · 被引用 243 次
- Hierarchical Memory Matching Network for Video Object SegmentationHongje Seong, Seoung Wug Oh, Joon-Young Lee, Seongwon Lee 等ICCV 2021 · 被引用 126 次
相关 Paper
- Learning Object State Changes in Videos: An Open-World PerspectiveZihui Xue, Kumar Ashutosh, Kristen GraumanCVPR 2024 · 被引用 12 次
- Breaking the "Object" in Video Object SegmentationPavel Tokmakov, Jie Li, Adrien GaidonCVPR 2023
- Unidentified Video Objects: A Benchmark for Dense, Open-World SegmentationWeiyao Wang, Matt Feiszli, Heng Wang, Du TranICCV 2021 · 被引用 151 次
- Segment Anything Across Shots: A Method and BenchmarkHengrui Hu, Kaining Ying, Henghui DingAAAI 2026 · 被引用 1 次
- MOSCATO: Predicting Multiple Object State Change through ActionsParnian Zameni, Yuhan Shen, Ehsan ElhamifarICCV 2025 · 被引用 4 次
