Depth-Aware Test-Time Training for Zero-Shot Video Object Segmentation
Weihuang Liu, Xi Shen, Haolun Li, Xiuli Bi, Bo Liu, Chi-Man Pun, Xiaodong Cun
摘要
Zero-shot Video Object Segmentation (ZSVOS) aims at segmenting the primary moving object without any human annotations. Mainstream solutions mainly focus on learning a single model on large-scale video datasets, which struggle to generalize to unseen videos. In this work, we introduce a test-time training (TTT) strategy to address the problem. Our key insight is to enforce the model to predict consistent depth during the TTT process. In detail, we first train a single network to perform both segmentation and depth prediction tasks. This can be effectively learned with our specifically designed depth modulation layer. Then, for the TTT process, the model is updated by predicting consistent depth maps for the same frame under different data augmentations. In addition, we explore different TTT weight updating strategies. Our empirical results suggest that the momentum-based weight initialization and looping-based training scheme lead to more stable improvements. Experiments show that the proposed method achieves clear improvements on ZSVOS. Our proposed video TTT strategy provides significant superiority over state-of-the-art TTT methods. Our code is available at: https://nifangbaage.github.io/DATTT/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Beyond Entropy: Region Confidence Proxy for Wild Test-Time AdaptationZixuan Hu, Yichun Hu, Xiaotong Li, Shixiang Tang 等ICML 2025
- Order-aware Interactive SegmentationBin Wang, Anwesa Choudhuri, Meng Zheng, Zhongpai Gao 等ICLR 2025
- Adaptive Dual Uncertainty Optimization: Boosting Monocular 3D Object Detection under Test-Time ShiftsZixuan Hu, Dongxiao Li, Xinzhu Ma, Shixiang Tang 等ICCV 2025
- Beyond Appearance: Camouflaged Object Detection via Geometric StructureJinyu Han, Changguang Wu, Fuming Sun, Jinhui TangCVPR 2026
- Shallow Features Matter: Hierarchical Memory with Heterogeneous Interaction for Unsupervised Video Object SegmentationXiangyu Zheng, Songcheng He, Wanyun Li, Xiaoqiang Li 等ACM MM 2025
它引用的顶会 Paper24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 被引用 2,416 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
相关 Paper
- Test-time Training for Matching-based Video Object SegmentationJuliette Bertrand, Giorgos Kordopatis-Zilos, Yannis Kalantidis, Giorgos ToliasNeurIPS 2023 · 被引用 7 次
- Consistent depth of moving objects in videoZhoutong Zhang, Forrester Cole, Richard Tucker, William T. Freeman 等SIGGRAPH 2021 · 被引用 26 次
- TTT++: When Does Self-Supervised Test-Time Training Fail or Thrive?Yuejiang Liu, Parth Kothari, Bastien van Delft, Baptiste Bellot-Gurlet 等NeurIPS 2021 · 被引用 469 次
- NC-TTT: A Noise Constrastive Approach for Test-Time TrainingDavid Osowiechi, Gustavo Adolfo Vargas Hakim, Mehrdad Noori, Milad Cheraghalikhani 等CVPR 2024 · 被引用 10 次
- Motion-Attentive Transition for Zero-Shot Video Object SegmentationTianfei Zhou, Shunzhou Wang, Yi Zhou, Yazhou Yao 等AAAI 2020 · 被引用 210 次
