SimulFlow: Simultaneously Extracting Feature and Identifying Target for Unsupervised Video Object Segmentation
Lingyi Hong, Wei Zhang, Shuyong Gao, Hong Lu, Wenqiang Zhang
Abstract
Unsupervised video object segmentation (UVOS) aims at detecting the primary objects in a given video sequence without any human interposing. Most existing methods rely on two-stream architectures that separately encode the appearance and motion information before fusing them to identify the target and generate object masks. However, this pipeline is computationally expensive and can lead to suboptimal performance due to the difficulty of fusing the two modalities properly. In this paper, we propose a novel UVOS model called SimulFlow that simultaneously performs feature extraction and target identification, enabling efficient and effective unsupervised video object segmentation. Concretely, we design a novel SimulFlow Attention mechanism to bridege the image and motion by utilizing the flexibility of attention operation, where coarse masks predicted from fused feature at each stage are used to constrain the attention operation within the mask area and exclude the impact of noise. Because of the bidirectional information flow between visual and optical flow features in SimulFlow Attention, no extra hand-designed fusing module is required and we only adopt a light decoder to obtain the final prediction. We evaluate our method on several benchmark datasets and achieve state-of-the-art results. Our proposed approach not only outperforms existing methods but also addresses the computational complexity and fusion difficulties caused by two-stream architectures. Our models achieve 87.4% J &F on DAVIS-16 with the highest speed (63.7 FPS on a 3090) and the lowest parameters (13.7 M). Our SimulFlow also obtains competitive results on video salient object detection datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f0eace3-b84e-4eaf-8fa6-0694f57e205bCited by top-tier papers5
- X-Prompt: Multi-modal Visual Prompt for Video Object SegmentationPinxue Guo, Wanyun Li, Hao Huang, Lingyi Hong et al.ACM MM 2024 · 7 citations
- General Compression Framework for Efficient Transformer Object TrackingLingyi Hong, Jinglun Li, Xinyu Zhou, Shilin Yan et al.ICCV 2025 · 5 citations
- Scoring, Remember, and Reference: Catching Camouflaged Objects in VideosYu'ang Feng, Shuyong Gao, Fuzhen Yan, Yicheng Song et al.ICCV 2025 · 2 citations
- Shallow Features Matter: Hierarchical Memory with Heterogeneous Interaction for Unsupervised Video Object SegmentationXiangyu Zheng, Songcheng He, Wanyun Li, Xiaoqiang Li et al.ACM MM 2025
- OneTracker: Unifying Visual Object Tracking with Foundation Models and Efficient TuningLingyi Hong, Shilin Yan, Renrui Zhang, Wanyun Li et al.CVPR 2024
Builds on26
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 746 citations
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 403 citations
Related papers
- Dual Prototype Attention for Unsupervised Video Object SegmentationSuhwan Cho, Minhyeok Lee, Seunghoon Lee, Dogyoon Lee et al.CVPR 2024
- Motion Guided Attention for Video Salient Object DetectionHaofeng Li, Guanqi Chen, Guanbin Li, Yizhou YuICCV 2019 · 200 citations
- Scalable Video Object Segmentation with Simplified FrameworkQiangqiang Wu, Tianyu Yang, Wei Wu, Antoni B. ChanICCV 2023 · 48 citations
- Bootstrapping Objectness from Videos by Relaxed Common Fate and Visual GroupingLong Lian, Zhirong Wu, Stella X. YuCVPR 2023
- VONet: Unsupervised Video Object Learning With Parallel U-Net Attention and Object-wise Sequential VAEHaonan Yu, Wei XuICLR 2024 · 1 citation
