A Filtering Framework for Semi-online Referring Video Object Segmentation
Xiao Hu, Heiko Neumann, Jochen Lang
Abstract
Referring video object segmentation (RVOS) extracts objects from videos based on provided text narrations. Previous approaches typically work on all video frames simultaneously through offline processing, but this is not always possible. Offline processing becomes also ineffective for long videos, as directly cutting the video into short clips to fit memory limits results in the loss of temporal consistency. To make RVOS more applicable to real-world video streams or long video scenarios, we introduce a Filtering Framework for RVOS (FF-RVOS), the first model capable of operating in online, semi-online, and offline modes with just one training session. We redesign RVOS as a stochastic optimization problem and leverage filtering to optimize object states across temporal sequences. Our method enhances temporal consistency for video streams and when a long video is processed in shorter clip sequences due to memory limitations. FF-RVOS demonstrates superior performance compared to previous state-of-the-art methods on public benchmarks with clear improvements, especially when the video is cut into short clips for processing. Our framework can also be embedded into different offline methods to boost temporal consistency. The project page is https://github.com/haliphinx/FF-RVOS.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 492f4ff3-2271-49ae-b844-7b1af711c858Related papers
- LoSh: Long-Short Text Joint Prediction Network for Referring Video Object SegmentationLinfeng Yuan, Miaojing Shi, Zijie Yue, Qijun ChenCVPR 2024 · 12 citations
- Tracking-forced Referring Video Object SegmentationRuxue Yan, Wenya Guo, Xubo Liu, Xumeng Liu et al.ACM MM 2024 · 3 citations
- Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object SegmentationTianming Liang, Haichao Jiang, Yuting Yang, Chaolei Tan et al.CVPR 2026 · 8 citations
- OnlineRefer: A Simple Online Baseline for Referring Video Object SegmentationDongming Wu, Tiancai Wang, Yuang Zhang, Xiangyu Zhang et al.ICCV 2023 · 82 citations
- Towards Streaming Referring Video Segmentation via Large Language ModelWenkang Zhang, Kaicheng Yang, Xiang An, Qiang Li et al.CVPR 2026
