Towards Robust Video Object Segmentation with Adaptive Object Calibration
Xiaohao Xu, Jinglu Wang, Xiang Ming, Yan Lu
Abstract
In the booming video era, video segmentation attracts increasing research attention in the multimedia community. Semi-supervised video object segmentation (VOS) aims at segmenting objects in all target frames of a video, given annotated object masks of reference frames. Most existing methods build pixel-wise reference-target correlations and then perform pixel-wise tracking to obtain target masks. Due to neglecting object-level cues, pixel-level approaches make the tracking vulnerable to perturbations, and even indiscriminate among similar objects. Towards robust VOS, the key insight is to calibrate the representation and mask of each specific object to be expressive and discriminative. Accordingly, we propose a new deep network, which can adaptively construct object representations and calibrate object masks to achieve stronger robustness. First, we construct the object representations by applying an adaptive object proxy (AOP) aggregation method, where the proxies represent arbitrary-shaped segments at multi-levels for reference. Then, prototype masks are initially generated from the reference-target correlations based on AOP. Afterwards, such proto-masks are further calibrated through network modulation, conditioning on the object proxy representations. We consolidate this conditional mask calibration process in a progressive manner, where the object representations and proto-masks evolve to be discriminative iteratively. Extensive experiments are conducted on the standard VOS benchmarks, YouTube-VOS-18/19 and DAVIS-17. Our model achieves the state-of-the-art performance among existing published works, and also exhibits superior robustness against perturbations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Where's Waldo: Diffusion Features For Personalized Segmentation and RetrievalDvir Samuel, Rami Ben-Ari, Matan Levy, Nir Darshan et al.NeurIPS 2024 · 19 citations
- Alignment Before Aggregation: Trajectory Memory Retrieval Network for Video Object SegmentationRui Sun, Yuan Wang, Huayu Mai, Tianzhu Zhang et al.ICCV 2023 · 12 citations
- Putting the Object Back into Video Object SegmentationHo Kei Cheng, Seoung Wug Oh, Brian L. Price, Joon-Young Lee et al.CVPR 2024
- Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any GranularityHuaxin Zhang, Xiaohao Xu, Xiang Wang, Jialong Zuo et al.CVPR 2025
- Scalable Benchmarking and Robust Learning for Noise-Free Ego-Motion and 3D Reconstruction from Noisy VideoXiaohao Xu, Tianyi Zhang, Shibo Zhao, Xiang Li et al.ICLR 2025
Builds on22
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 403 citations
- Associating Objects with Transformers for Video Object SegmentationZongxin Yang, Yunchao Wei, Yi YangNeurIPS 2021 · 398 citations
- MBRS: Enhancing Robustness of DNN-based Watermarking by Mini-Batch of Real and Simulated JPEG CompressionZhaoyang Jia, Han Fang, Weiming ZhangACM MM 2021 · 251 citations
Related papers
- Per-Clip Video Object SegmentationKwanyong Park, Sanghyun Woo, Seoung Wug Oh, In So Kweon et al.CVPR 2022 · 45 citations
- Dual Temporal Memory Network for Efficient Video Object SegmentationKaihua Zhang, Long Wang, Dong Liu, Bo Liu et al.ACM MM 2020 · 16 citations
- Video Object Segmentation with Dynamic Memory Networks and Adaptive Object AlignmentShuxian Liang, Xu Shen, Jianqiang Huang, Xian-Sheng HuaICCV 2021 · 28 citations
- Reliable Propagation-Correction Modulation for Video Object SegmentationXiaohao Xu, Jinglu Wang, Xiao Li, Yan LuAAAI 2022 · 74 citations
- Spatiotemporal Graph Neural Network based Mask Reconstruction for Video Object SegmentationDaizong Liu, Shuangjie Xu, Xiao-Yang Liu, Zichuan Xu et al.AAAI 2021 · 25 citations
