Spatial-temporal Causal Inference for Partial Image-to-video Adaptation
Jin Chen, Xinxiao Wu, Yao Hu, Jiebo Luo
Abstract
Image-to-video adaptation leverages off-the-shelf learned models in labeled images to help classification in unlabeled videos, thus alleviating the high computation overhead of training a video classifier from scratch. This task is very challenging since there exist two types of domain shifts between images and videos: 1) spatial domain shift caused by static appearance variance between images and video frames, and 2) temporal domain shift caused by the absence of dynamic motion in images. Moreover, for different video classes, these two domain shifts have different effects on the domain gap and should not be treated equally during adaptation. In this paper, we propose a spatial-temporal causal inference framework for image-to-video adaptation. We first construct a spatial-temporal causal graph to infer the effects of the spatial and temporal domain shifts by performing counterfactual causality. We then learn causality-guided bidirectional heterogeneous mappings between images and videos to adaptively reduce the two domain shifts. Moreover, to relax the assumption that the label spaces of the image and video domains are the same by the existing methods, we incorporate class-wise alignment into the learning of image-video mappings to perform partial image-to-video adaptation where the image label space subsumes the video label space. Extensive experiments on several video datasets have validated the effectiveness of our proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc73e586-e251-4a91-97e9-c87747082286Cited by top-tier papers3
- Robust Event Forecasting with Spatiotemporal Confounder LearningSonggaojun Deng, Huzefa Rangwala, Yue NingKDD 2022 · 9 citations
- Composing Concepts from Images and Videos via Concept-prompt BindingXianghao Kong, Zeyu Zhang, Yuwei Guo, Zhuoran Zhao et al.CVPR 2026 · 2 citations
- Image-to-video Adaptation with Outlier Modeling and Robust Self-learningJunbao Zhuo, Shuhui Wang, Zhenghan Chen, Li Shen et al.AAAI 2025
Builds on6
- Larger Norm More Transferable: An Adaptive Feature Norm Approach for Unsupervised Domain AdaptationRuijia Xu, Guanbin Li, Jihan Yang, Liang LinICCV 2019 · 563 citations
- Counterfactual Critic Multi-Agent Training for Scene Graph GenerationLong Chen, Hanwang Zhang, Jun Xiao, Xiangnan He et al.ICCV 2019 · 165 citations
- Visual Commonsense R-CNNTan Wang, Jianqiang Huang, Hanwang Zhang, Qianru SunCVPR 2020
- Counterfactual Samples Synthesizing for Robust Visual Question AnsweringLong Chen, Xin Yan, Jun Xiao, Hanwang Zhang et al.CVPR 2020
- Universal Source-Free Domain AdaptationJogendra Nath Kundu, Naveen Venkat, Rahul M. V., R. Venkatesh BabuCVPR 2020
Related papers
- Synthesizing Videos from Images for Image-to-Video AdaptationJunbao Zhuo, Xingyu Zhao, Shuhui Wang, Huimin Ma et al.ACM MM 2023 · 4 citations
- Adversarial Bipartite Graph Learning for Video Domain AdaptationYadan Luo, Zi Huang, Zijian Wang, Zheng Zhang et al.ACM MM 2020 · 40 citations
- Relative Alignment Network for Source-Free Multimodal Video Domain AdaptationYi Huang, Xiaoshan Yang, Ji Zhang, Changsheng XuACM MM 2022 · 18 citations
- Adaptive Image-to-Video Scene Graph Generation via Knowledge Reasoning and Adversarial LearningJin Chen, Xiaofeng Ji, Xinxiao WuAAAI 2022 · 3 citations
- Partial Video Domain Adaptation with Partial Adversarial Temporal Attentive NetworkYuecong Xu, Jianfei Yang, Haozhi Cao, Zhenghua Chen et al.ICCV 2021 · 32 citations
