Modality-Collaborative Test-Time Adaptation for Action Recognition
Baochen Xiong, Xiaoshan Yang, Yaguang Song, Yaowei Wang, Changsheng Xu
摘要
Video-based Unsupervised Domain Adaptation (VUDA) method improves the generalization of the video model, en-abling it to be applied to action recognition tasks in different environments. However, these methods require contin-uous access to source data during the adaptation process, which are impractical in real scenarios where the source videos are not available with concerns in transmission efficiency or privacy issues. To address this problem, in this paper, we focus on the Multimodal Video Test- Time Adaptation (MVTTA) task. Existing image-based TTA methods cannot be directly applied to this task because videos have domain shifts in multimodal and temporal, which brings difficulties to adaptation. To address the above challenges, we propose a Modality-Collaborative Test-Time Adaptation (MC-TTA) Network. MC-TTA contains maintain teacher and student memory banks respectively for generating pseudo-prototypes and target-prototypes. In the teacher model, we propose Self-assembled Source-friendly Feature Reconstruction (SSFR) to encourage the teacher memory bank to store features that are more likely to be consistent with the source distribution. Through multimodal prototype alignment and cross-modal relative consistency, our method can effectively alleviate domain shift in videos. We evaluate the proposed model on four public video datasets. The results show that our model outperforms existing state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- HiVG: Hierarchical Multimodal Fine-grained Modulation for Visual GroundingLinhui Xiao, Xiaoshan Yang, Fang Peng, Yaowei Wang 等ACM MM 2024 · 被引用 28 次
- Pilot: Building the Federated Multimodal Instruction Tuning FrameworkBaochen Xiong, Xiaoshan Yang, Yaguang Song, Yaowei Wang 等AAAI 2025 · 被引用 6 次
- Adversarial Alignment with Anchor Dragging Drift (A³D²): Multimodal Domain Adaptation with Partially Shifted ModalitiesJun Sun, Xinxin Zhang, Simin Hong, Jian Zhu 等ACL 2025 · 被引用 5 次
- Bilateral Collaboration with Large Vision-Language Models for Open Vocabulary Human-Object Interaction DetectionYupeng Hu, Changxing Ding, Chang Sun, Shaoli Huang 等ICCV 2025 · 被引用 1 次
- Analytic Continual Test-Time Adaptation for Multi-Modality CorruptionYufei Zhang, Yicheng Xu, Hongxin Wei, Zhiping Lin 等ACM MM 2025 · 被引用 1 次
它引用的顶会 Paper27
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
- Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain AdaptationJian Liang, Dapeng Hu, Jiashi FengICML 2020 · 被引用 1,624 次
- Efficient Test-Time Model Adaptation without ForgettingShuaicheng Niu, Jiaxiang Wu, Yifan Zhang, Yaofo Chen 等ICML 2022 · 被引用 579 次
相关 Paper
- Relative Alignment Network for Source-Free Multimodal Video Domain AdaptationYi Huang, Xiaoshan Yang, Ji Zhang, Changsheng XuACM MM 2022 · 被引用 18 次
- Dynamic-Static Collaboration for Unsupervised Domain Adaptive Video-Based Visible-Infrared Person Re-IdentificationJiaxu Leng, Zhengjie Wang, Shuang Li, Xinbo GaoAAAI 2026
- Bridging Modalities via Progressive Re-alignment for Multimodal Test-Time AdaptationJiacheng Li, Songhe FengAAAI 2026 · 被引用 2 次
- Unsupervised Video Domain Adaptation with Masked Pre-Training and Collaborative Self-TrainingArun V. Reddy, William Paul, Corban Rivera, Ketul Shah 等CVPR 2024 · 被引用 3 次
- Mix-DANN and Dynamic-Modal-Distillation for Video Domain AdaptationYuehao Yin, Bin Zhu, Jingjing Chen, Lechao Cheng 等ACM MM 2022 · 被引用 7 次
