Generalizable Fourier Augmentation for Unsupervised Video Object Segmentation
Huihui Song, Tiankang Su, Yuhui Zheng, Kaihua Zhang, Bo Liu, Dong Liu
Abstract
The performance of existing unsupervised video object segmentation methods typically suffers from severe performance degradation on test videos when tested in out-of-distribution scenarios. The primary reason is that the test data in realworld may not follow the independent and identically distribution (i.i.d.) assumption, leading to domain shift. In this paper, we propose a Generalizable Fourier Augmentation (G-FA) method during training to improve the generalization ability of the model. To achieve this, the GFA performs Fast Fourier Transform (FFT) over the intermediate spatial domain features in each layer to yield corresponding frequency representations, including amplitude components (encoding scene-aware styles such as texture, color, contrast of the scene) and phase components (encoding rich semantics). We produce a variety of style features via Gaussian sampling to augment the training data, thereby improving the generalization capability of the model. To further improve the crossdomain generalization performance of the model, we design a phase feature update strategy via exponential moving average using phase features from past frames in an online update manner, which could help the model to learn cross-domaininvariant features. Extensive experiments show that the proposed GFA achieves the state-of-the-art performance on popular benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- FIM: Frequency-Aware Multi-View Interest Modeling for Local-Life Service RecommendationGuoquan Wang, Qiang Luo, Weisong Hu, Pengfei Yao et al.SIGIR 2025 · 7 citations
- Diffusion-Based Source-Biased Model for Single Domain Generalized Object DetectionHan Jiang, Wenfei Yang, Tianzhu Zhang, Yongdong ZhangICCV 2025 · 2 citations
- Shallow Features Matter: Hierarchical Memory with Heterogeneous Interaction for Unsupervised Video Object SegmentationXiangyu Zheng, Songcheng He, Wanyun Li, Xiaoqiang Li et al.ACM MM 2025
Builds on10
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Uncertainty Modeling for Out-of-Distribution GeneralizationXiaotong Li, Yongxing Dai, Yixiao Ge, Jun Liu et al.ICLR 2022 · 237 citations
- Motion-Attentive Transition for Zero-Shot Video Object SegmentationTianfei Zhou, Shunzhou Wang, Yi Zhou, Yazhou Yao et al.AAAI 2020 · 210 citations
- Full-Duplex Strategy for Video Object SegmentationGe-Peng Ji, Keren Fu, Zhe Wu, Deng-Ping Fan et al.ICCV 2021 · 173 citations
- Feature Stylization and Domain-aware Contrastive Learning for Domain GeneralizationSeogkyu Jeon, Kibeom Hong, Pilhyeon Lee, Jewook Lee et al.ACM MM 2021 · 77 citations
Related papers
- Unsupervised Video Object Segmentation with Online Adversarial Self-TuningTiankang Su, Huihui Song, Dong Liu, Bo Liu et al.ICCV 2023 · 19 citations
- A Fourier-Based Framework for Domain GeneralizationQinwei Xu, Ruipeng Zhang, Ya Zhang, Yanfeng Wang et al.CVPR 2021
- Generalizable Multi-Camera 3D Object Detection from a Single Source via Fourier Cross-View LearningXue Zhao, Qinying Gu, Xinbing Wang, Chenghu Zhou et al.ICML 2025
- Consistency Learning based on Class-Aware Style Variation for Domain Generalizable Semantic SegmentationSiwei Su, Haijian Wang, Meng YangACM MM 2022 · 10 citations
- Adversarial Style Augmentation for Domain Generalized Urban-Scene SegmentationZhun Zhong, Yuyang Zhao, Gim Hee Lee, Nicu SebeNeurIPS 2022 · 130 citations
