Image-to-video Adaptation with Outlier Modeling and Robust Self-learning
Junbao Zhuo, Shuhui Wang, Zhenghan Chen, Li Shen, Qingming Huang, Huimin Ma
Abstract
The image-to-video adaptation task seeks to effectively harness both labeled images and unlabeled videos for achieving effective video recognition. The modality gap of the image and video modalities and the domain discrepancy across the two domains are the two essential challenges in this task. Existing methods reduce the domain discrepancy via close-set domain adaptation techniques, resulting in inaccurate domain alignment as there exist outlier target frames. To tackle this issue, we extend the vanilla classifier with outlier classes, where each outlier class responsible for capturing outlier frames for a specific class via batch nuclear norm maximization loss. We further propose a new loss by treating the source images apart from class c as instances from outlier class specific for c. As for the modality gap, existing methods usually utilize the pseudo labels obtained from an image-level adapted model to learn a video-level model. Rare efforts are dedicated to handling the noise in pseudo labels. We proposed a new metric based on label propagation consistency to select samples for training a better video-level model. Experiments on 3 benchmarks validating the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e304ef17-c725-42b2-b2c0-0ec439842272Builds on8
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin et al.CVPR 2022 · 752 citations
- Theoretical Analysis of Self-Training with Deep Networks on Unlabeled DataColin Wei, Kendrick Shen, Yining Chen, Tengyu MaICLR 2021 · 261 citations
- OpenMatch: Open-Set Semi-supervised Learning with Open-set Consistency RegularizationKuniaki Saito, Donghyun Kim, Kate SaenkoNeurIPS 2021 · 80 citations
Related papers
- Synthesizing Videos from Images for Image-to-Video AdaptationJunbao Zhuo, Xingyu Zhao, Shuhui Wang, Huimin Ma et al.ACM MM 2023 · 4 citations
- Spatial-temporal Causal Inference for Partial Image-to-video AdaptationJin Chen, Xinxiao Wu, Yao Hu, Jiebo LuoAAAI 2021 · 19 citations
- Calibrating Class Weights with Multi-Modal Information for Partial Video Domain AdaptationXiyu Wang, Yuecong Xu, Jianfei Yang, Kezhi MaoACM MM 2022 · 6 citations
- Relative Alignment Network for Source-Free Multimodal Video Domain AdaptationYi Huang, Xiaoshan Yang, Ji Zhang, Changsheng XuACM MM 2022 · 18 citations
- Tell, Don't Show: Language Guidance Eases Transfer Across Domains in Images and VideosTarun Kalluri, Bodhisattwa Prasad Majumder, Manmohan ChandrakerICML 2024 · 7 citations
