Synthesizing Videos from Images for Image-to-Video Adaptation
Junbao Zhuo, Xingyu Zhao, Shuhui Wang, Huimin Ma, Qingming Huang
Abstract
We address the image-to-video adaptation task that aims to leverage labeled images and unlabeled videos for video recognition. There are two major challenges in this task, including the domain discrepancy between the two domains, and the modality gap between the image and video modalities. Existing methods mainly employ a two-stage paradigm by first adopting frame-level adaptation to reduce the domain discrepancy and then learning a spatio-temporal model to bridge the modality gap. In this paper, we provide a new perspective and propose a single-stage method that synthesizes video from the source static image and converts the image-to-video adaptation problem into a video-to-video adaptation problem. With the synthesized video, we present a simple baseline that a spatio-temporal model is trained with cross entropy loss with source labels and the Batch Nuclear norm Maximization loss to encourage the classification responses of target videos maintain the discriminability and diversity. We further propose a new pseudo label generation method that inherits the robustness of class prototype and the effectiveness of the small loss criterion. Based on the constructed baseline and the proposed pseudo label generation method, we train a model that achieves state-of-the-art performances or gets comparable performances on three standard benchmarks. Our codes are publicly available at https://github.com/junbaoZHUO/ST-I2V.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 7557a068-0d50-4b61-9df1-83c9fc8a340cCited by top-tier papers3
- A Recipe for Scaling up Text-to-Video Generation with Text-free VideosXiang Wang, Shiwei Zhang, Hangjie Yuan, Zhiwu Qing et al.CVPR 2024 · 19 citations
- Zero-Shot Controllable Image-to-Video Animation via Motion DecompositionShoubin Yu, Jacob Zhiyuan Fang, Jian Zheng, Gunnar A. Sigurdsson et al.ACM MM 2024 · 3 citations
- EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX GenerationYue Ma, Xu Ye, Qinghe Wang, Yucheng Wang et al.SIGGRAPH 2026 · 2 citations
Related papers
- Image-to-video Adaptation with Outlier Modeling and Robust Self-learningJunbao Zhuo, Shuhui Wang, Zhenghan Chen, Li Shen et al.AAAI 2025
- Unsupervised Image-to-Video Adaptation via Category-aware Flow Memory Bank and Realistic Video GenerationKenan Huang, Junbao Zhuo, Shuhui Wang, Chi Su et al.ACM MM 2024
- Spatial-temporal Causal Inference for Partial Image-to-video AdaptationJin Chen, Xinxiao Wu, Yao Hu, Jiebo LuoAAAI 2021 · 19 citations
- Relative Alignment Network for Source-Free Multimodal Video Domain AdaptationYi Huang, Xiaoshan Yang, Ji Zhang, Changsheng XuACM MM 2022 · 18 citations
- Unsupervised Video Domain Adaptation with Masked Pre-Training and Collaborative Self-TrainingArun V. Reddy, William Paul, Corban Rivera, Ketul Shah et al.CVPR 2024 · 3 citations
