Hierarchical Debiasing and Noisy Correction for Cross-domain Video Tube Retrieval
Jingqiao Xiu, Mengze Li, Wei Ji, Jingyuan Chen, Hanbin Zhao, Shin'ichi Satoh, Roger Zimmermann
Abstract
Video Tube Retrieval (VTR) has attracted wide attention in the multi-modal domain, aiming to accurately localize the spatial-temporal tube in videos based on the natural language description. Despite the remarkable progress, existing VTR models trained on a specific domain (source domain) often perform unsatisfactory in another domain (target domain), due to the domain gap. Toward this issue, we introduce the learning strategy, Unsupervised Domain Adaptation, into the VTR task (UDA-VTR), which enables the knowledge transfer from the labeled source domain to the unlabeled target domain without additional manual annotations. An intuitive solution is generating the pseudo labels for the target domain samples with the fully trained source model and fine-tuning the source model on the target domain with pseudo labels. However, the existing domain gap gives rise to two problems for this process: (1) The transfer of model parameters across domains may introduce source domain bias into target domain features, significantly impacting the feature-based prediction for target domain samples. (2) The pseudo labels tend to identify video tubes that are widely present in the source domain, rather than accurately localizing the correct video tubes specific to the target domain samples. To address the above issues, we propose the unsupervised domain adaptation model via Hierarchical dEbiAsing and noisy corRecTion (HEART) for cross-domain video tube retrieval, which contains two characteristic modules: Layered Feature Debiasing (including the adversarial feature alignment and the graph based alignment) and Pseudo Label Refinement. Extensive experiments prove the effectiveness of our HEART model by significantly surpassing the state-of-the-arts.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 4458e8a4-7e4f-4eee-a0aa-a6026edf56b7Cited by top-tier papers4
- Multi-Pair Temporal Sentence Grounding via Multi-Thread Knowledge Transfer NetworkXiang Fang, Wanlong Fang, Changshuo Wang, Daizong Liu et al.AAAI 2025 · 10 citations
- Few-Shot Incremental Learning via Foreground Aggregation and Knowledge Transfer for Audio-Visual Semantic SegmentationJingqiao Xiu, Mengze Li, Zongxin Yang, Wei Ji et al.AAAI 2025 · 3 citations
- Geometric Alignment and Prior Modulation for View-Guided Point Cloud Completion on Unseen CategoriesJingqiao Xiu, Yicong Li, Na Zhao, Han Fang et al.ICCV 2025 · 2 citations
- MLLMSplat: A 2D MLLM-Powered Framework for 3D Gaussian Splatting Understanding, Generation, and EditingJingqiao Xiu, Can Wang, Dong XuCVPR 2026
Related papers
- Unsupervised Domain Adaptative Temporal Sentence Localization with Mutual Information MaximizationDaizong Liu, Xiang Fang, Xiaoye Qu, Jianfeng Dong et al.AAAI 2024 · 9 citations
- Uncertainty-Aware Alignment Network for Cross-Domain Video-Text RetrievalXiaoshuai Hao, Wanqian ZhangNeurIPS 2023 · 26 citations
- Mind-the-Gap! Unsupervised Domain Adaptation for Text-Video RetrievalQingchao Chen, Yang Liu, Samuel AlbanieAAAI 2021 · 28 citations
- Relative Alignment Network for Source-Free Multimodal Video Domain AdaptationYi Huang, Xiaoshan Yang, Ji Zhang, Changsheng XuACM MM 2022 · 18 citations
- Unsupervised Domain Adaptation for Video Object Grounding with Cascaded Debiasing LearningMengze Li, Haoyu Zhang, Juncheng Li, Zhou Zhao et al.ACM MM 2023 · 7 citations
