Breaking the "Object" in Video Object Segmentation
Pavel Tokmakov, Jie Li, Adrien Gaidon
摘要
The appearance of an object can be fleeting when it transforms. As eggs are broken or paper is torn, their color, shape and texture can change dramatically, preserving virtually nothing of the original except for the identity itself. Yet, this important phenomenon is largely absent from existing video object segmentation (VOS) benchmarks. In this work, we close the gap by collecting a new dataset for Video Object Segmentation under Transformations (VOST). It consists of more than 700 high-resolution videos, captured in diverse environments, which are 20 seconds long on average and densely labeled with instance masks. We adopt a careful, multi-step approach to ensure that these videos focus on complex object transformations, capturing their full temporal extent. We then extensively evaluate stateof-the-art VOS methods and make a number of important discoveries. In particular, we show that existing methods struggle when applied to this novel task and that their main limitation lies in over-reliance on static appearance cues. This motivates us to propose a few modifications for the topperforming baseline that improve its capabilities by better modeling spatio-temporal information. More broadly, our work highlights the need for further research on learning more robust video object representations. Nothing is lost or created, all things are merely transformed.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- RMem: Restricted Memory Banks Improve Video Object SegmentationJunbao Zhou, Ziqi Pang, Yu-Xiong WangCVPR 2024 · 被引用 18 次
- Video State-Changing Object SegmentationJiangwei Yu, Xiang Li, Xinran Zhao, Hongming Zhang 等ICCV 2023 · 被引用 16 次
- SAM2LONG: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory TreeShuangrui Ding, Rui Qian, Xiaoyi Dong, Pan Zhang 等ICCV 2025 · 被引用 15 次
- Learning Object State Changes in Videos: An Open-World PerspectiveZihui Xue, Kumar Ashutosh, Kristen GraumanCVPR 2024 · 被引用 12 次
- DynamicVerse: A Physically-Aware Multimodal Framework for 4D World ModelingKairun Wen, Yuzhi Huang, Runyu Chen, Hui Zheng 等NeurIPS 2025 · 被引用 11 次
它引用的顶会 Paper20
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 被引用 845 次
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
相关 Paper
- Tracking and Understanding Object TransformationsYihong Sun, Xinyu Yang, Jennifer J. Sun, Bharath HariharanNeurIPS 2025 · 被引用 5 次
- MOSE: A New Dataset for Video Object Segmentation in Complex ScenesHenghui Ding, Chang Liu, Shuting He, Xudong Jiang 等ICCV 2023 · 被引用 267 次
- Unidentified Video Objects: A Benchmark for Dense, Open-World SegmentationWeiyao Wang, Matt Feiszli, Heng Wang, Du TranICCV 2021 · 被引用 151 次
- M^3-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object SegmentationZixuan Chen, Jiaxin Li, Junxuan Liang, Liming Tan 等CVPR 2025
- Audio-Visual Instance SegmentationRuohao Guo, Xianghua Ying, Yaru Chen, Dantong Niu 等CVPR 2025
