YouMVOS: An Actor-centric Multi-shot Video Object Segmentation Dataset
Donglai Wei, Siddhant Kharbanda, Sarthak Arora, Roshan Roy, Nishant Jain, Akash Palrecha, Tanav Shah, Shray Mathur, Ritik Mathur, Abhijay Kemkar, Anirudh Srinivasan Chakravarthy, Zudi Lin
Abstract
Many video understanding tasks require analyzing multishot videos, but existing datasets for video object segmentation (VOS) only consider single-shot videos. To address this challenge, we collected a new dataset-YouMVaS-of 200 popular YouTube videos spanning ten genres, where each video is on average five minutes long and with 75 shots. We selected recurring actors and annotated 431K segmentation masks at a frame rate of six, exceeding previous datasets in average video duration, object variation, and narrative structure complexity. We incorporated good practices of model architecture design, memory management, and multi-shot tracking into an existing video segmentation method to build competitive baseline methods. Through error analysis, we found that these baselines still fail to cope with cross-shot appearance variation on our YouMVOS dataset. Thus, our dataset poses new challenges in multi-shot segmentation towards better video analysis. Data, code, and pre-trained models are available at https://donglaiw.github.io/proj/youMVOS
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- AVTrack: Audio-Visual Tracking in Human-centric Complex ScenesYaoting Wang, Yun Zhou, Zipei Zhang, Henghui DingICML 2026 · 1 citation
- Segment Anything Across Shots: A Method and BenchmarkHengrui Hu, Kaining Ying, Henghui DingAAAI 2026 · 1 citation
Builds on14
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- ABD-Net: Attentive but Diverse Person Re-IdentificationTianlong Chen, Shaojin Ding, Jingyi Xie, Ye Yuan et al.ICCV 2019 · 544 citations
- Motion-Attentive Transition for Zero-Shot Video Object SegmentationTianfei Zhou, Shunzhou Wang, Yi Zhou, Yazhou Yao et al.AAAI 2020 · 210 citations
- Anchor Diffusion for Unsupervised Video Object SegmentationZhao Yang, Qiang Wang, Luca Bertinetto, Song Bai et al.ICCV 2019 · 127 citations
Related papers
- Unidentified Video Objects: A Benchmark for Dense, Open-World SegmentationWeiyao Wang, Matt Feiszli, Heng Wang, Du TranICCV 2021 · 151 citations
- LVOS: A Benchmark for Long-term Video Object SegmentationLingyi Hong, Wenchao Chen, Zhongying Liu, Wei Zhang et al.ICCV 2023 · 89 citations
- Multi-Granularity Video Object SegmentationSangbeom Lim, Seongchan Kim, Seungjun An, Seokju Cho et al.AAAI 2025
- Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object SegmentationTianming Liang, Haichao Jiang, Yuting Yang, Chaolei Tan et al.CVPR 2026 · 8 citations
- MOSE: A New Dataset for Video Object Segmentation in Complex ScenesHenghui Ding, Chang Liu, Shuting He, Xudong Jiang et al.ICCV 2023 · 267 citations
