Unsupervised Domain Adaptation for Video Object Grounding with Cascaded Debiasing Learning
Mengze Li, Haoyu Zhang, Juncheng Li, Zhou Zhao, Wenqiao Zhang, Shengyu Zhang, Shiliang Pu, Yueting Zhuang, Fei Wu
Abstract
This paper addresses the Unsupervised Domain Adaptation (UDA) for the dense frame prediction task - Video Object Grounding (VOG). This investigation springs from the recognition of the limited generalization capabilities of data-driven approaches when confronted with unseen test scenarios. We set the goal of enhancing the adaptability of the source-dominated model from a labeled domain to the unlabeled target domain through re-training on pseudo-labels (i.e., predicted boxes of language-described objects). Given the potential for source-domain biases in the pseudo-label generation, we decompose the labeling refinement as two cascaded debiasing subroutines: (1) we develop a discarded training strategy to correct the Biased Proposal Selection by filtering out the examples with uncertain proposals selected from the proposal (candidate box) set. The identifier of these uncertain examples is the discordance between the predictions of the source-dominated model and those of a target-domain clustered classifier, which remains free from the source-domain bias. (2) With the refined proposals as a foundation, we measure Grounding Coordinate Offset based on the semantic distance of the model's prediction across domains, based on which we alleviate source-domain bias in the target model through adversarial learning. To verify the superiority of the proposed method, we collected two UDA-VOG datasets called I2O-VOG and R2M-VOG by manually dividing and combining the well-known VOG datasets. The extensive experiments on them show our model significantly outperforms SOTA methods by a large margin.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6c18878c-3d3c-4dcd-aa22-b6b361e6e5feCited by top-tier papers5
- Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative InstructionsJuncheng Li, Kaihang Pan, Zhiqi Ge, Minghe Gao et al.ICLR 2024 · 95 citations
- A Unified Approach to Domain Incremental Learning with Memory: Theory and AlgorithmHaizhou Shi, Hao WangNeurIPS 2023 · 60 citations
- Revisiting the Domain Shift and Sample Uncertainty in Multi-source Active Domain TransferWenqiao Zhang, Zheqi LvCVPR 2024 · 17 citations
- Unsupervised Domain Adaptation for Anatomical Structure Detection in Ultrasound ImagesBin Pu, Xingguo Lv, Jiewen Yang, Guannan He et al.ICML 2024 · 10 citations
- Leveraging Anatomical Consistency for Multi-Object Detection in Ultrasound Images via Source-free Unsupervised Domain AdaptationBin Pu, Xingguo Lv, Jiewen Yang, Xingbo Dong et al.AAAI 2025 · 6 citations
Related papers
- Category Dictionary Guided Unsupervised Domain Adaptation for Object DetectionShuai Li, Jianqiang Huang, Xian-Sheng Hua, Lei ZhangAAAI 2021 · 47 citations
- Pseudo Label Refinery for Unsupervised Domain Adaptation on Cross-Dataset 3D Object DetectionZhanwei Zhang, Minghao Chen, Shuai Xiao, Liang Peng et al.CVPR 2024 · 10 citations
- Unsupervised Video Domain Adaptation with Masked Pre-Training and Collaborative Self-TrainingArun V. Reddy, William Paul, Corban Rivera, Ketul Shah et al.CVPR 2024 · 3 citations
- Hierarchical Debiasing and Noisy Correction for Cross-domain Video Tube RetrievalJingqiao Xiu, Mengze Li, Wei Ji, Jingyuan Chen et al.ACM MM 2024 · 5 citations
- IGG: Improved Graph Generation for Domain Adaptive Object DetectionPengteng Li, Ying He, F. Richard Yu, Pinhao Song et al.ACM MM 2023 · 10 citations
