LVOS: A Benchmark for Long-term Video Object Segmentation
Lingyi Hong, Wenchao Chen, Zhongying Liu, Wei Zhang, Pinxue Guo, Zhaoyu Chen, Wenqiang Zhang
Abstract
Existing video object segmentation (VOS) benchmarks focus on short-term videos which just last about 3-5 seconds and where objects are visible most of the time. These videos are poorly representative of practical applications, and the absence of long-term datasets restricts further investigation of VOS on the application in realistic scenarios. So, in this paper, we present a new benchmark dataset named LVOS, which consists of 220 videos with a total duration of 421 minutes. To the best of our knowledge, LVOS is the first densely annotated long-term VOS dataset. The videos in our LVOS last 1.59 minutes on average, which is 20 times longer than videos in existing VOS datasets. Each video includes various attributes, especially challenges deriving from the wild, such as long-term reappearing and cross-temporal similar objeccts. Based on LVOS, we assess existing video object segmentation algorithms and propose a Diverse Dynamic Memory network (DDMemory) that consists of three complementary memory banks to exploit temporal information adequately. The experimental results demonstrate the strength and weaknesses of prior methods, pointing promising directions for further study. Data and code are available at https://lingyihongfd.github.io/lvos.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 516ded07-59d4-4e72-9c56-a0b65e2e6faeCited by top-tier papers27
- Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and GroundingChristopher Clark, Jieyu Zhang, Zixian Ma, Jae Sung Park et al.CVPR 2026 · 144 citations
- XMem++: Production-level Video Segmentation From Few Annotated FramesMaksym Bekuzarov, Ariana Bermudez, Joon-Young Lee, Hao LiICCV 2023 · 69 citations
- Segment Every Reference Object in Spatial and Temporal SpacesJiannan Wu, Yi Jiang, Bin Yan, Huchuan Lu et al.ICCV 2023 · 29 citations
- RMem: Restricted Memory Banks Improve Video Object SegmentationJunbao Zhou, Ziqi Pang, Yu-Xiong WangCVPR 2024 · 18 citations
- Advancing Complex Video Object Segmentation via Progressive Concept ConstructionZhixiong Zhang, Shuangrui Ding, Xiaoyi Dong, Songxin He et al.ICLR 2026 · 17 citations
Builds on22
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 746 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 403 citations
- Associating Objects with Transformers for Video Object SegmentationZongxin Yang, Yunchao Wei, Yi YangNeurIPS 2021 · 398 citations
Related papers
- Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object SegmentationTianming Liang, Haichao Jiang, Yuting Yang, Chaolei Tan et al.CVPR 2026 · 8 citations
- Dual Temporal Memory Network for Efficient Video Object SegmentationKaihua Zhang, Long Wang, Dong Liu, Bo Liu et al.ACM MM 2020 · 16 citations
- LiVOS: Light Video Object Segmentation with Gated Linear MatchingQin Liu, Jianfeng Wang, Zhengyuan Yang, Linjie Li et al.CVPR 2025
- Video Object Segmentation with Dynamic Memory Networks and Adaptive Object AlignmentShuxian Liang, Xu Shen, Jianqiang Huang, Xian-Sheng HuaICCV 2021 · 28 citations
- Recurrent Dynamic Embedding for Video Object SegmentationMingxing Li, Li Hu, Zhiwei Xiong, Bang Zhang et al.CVPR 2022 · 80 citations
