Contrastive Disentangled Meta-Learning for Signer-Independent Sign Language Translation
Tao Jin, Zhou Zhao
摘要
Sign language translation aims at directly translating a sign language video into a natural sentence. The majority of existing methods take the video-sentence pairs labeled by multiple specific signers as training and testing samples. However, such setting does not fit in with the real-world applications. A practicable sign language translation system is supposed to provide accurate translation results for unseen signers. In this paper, we mainly attack the signer-independent setting and focus on augmenting the generalization ability of translation model. To adapt to the challenging setting, we propose a novel framework called contrastive disentangled meta-learning (CDM), which develops several improvements in both deep architecture and training mode. Specifically, based on the minimax entropy objective, a disentangled module with adaptive gated units is developed to decouple the signer-specific and task-specific representation in the encoder. Besides, we facilitate the frame-word alignments by leveraging contrastive constraints between the obtained task-specific representation and the decoding output. The disentangled and contrastive modules could provide complementary information for each other. As for the training mode, we encourage the model to perform well in the simulated signer-independent scenarios by finding the generalized learning directions in the meta-learning process. Considering that vanilla meta-learning methods utilize the multiple specific signers insufficiently, we adopt a fine-grained learning strategy that simultaneously conducts meta-learning in a variety of domain shift scenarios in each iteration. Extensive experiments on the benchmark dataset RWTH-PHOENIX-Weather-2014T(PHOENIX14T) show that CDM could achieve competitive results compared with the state-of-the-art methods.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper6
- BEST: BERT Pre-training for Sign Language Recognition with Coupling TokenizationWeichao Zhao, Hezhen Hu, Wengang Zhou, Jiaxin Shi 等AAAI 2023 · 被引用 70 次
- Improving Gloss-free Sign Language Translation by Reducing Representation DensityJinhui Ye, Xing Wang, Wenxiang Jiao, Junwei Liang 等NeurIPS 2024 · 被引用 49 次
- SimulSLT: End-to-End Simultaneous Sign Language TranslationAoxiong Yin, Zhou Zhao, Jinglin Liu, Weike Jin 等ACM MM 2021 · 被引用 35 次
- Exploring Group Video Captioning with Efficient Relational ApproximationWang Lin, Tao Jin, Ye Wang, Wenwen Pan 等ICCV 2023 · 被引用 17 次
- OpenSR: Open-Modality Speech Recognition via Maintaining Multi-Modality AlignmentXize Cheng, Tao Jin, Linjun Li, Wang Lin 等ACL 2023 · 被引用 10 次
相关 Paper
- MC-SLT: Towards Low-Resource Signer-Adaptive Sign Language TranslationTao Jin, Zhou Zhao, Meng Zhang, Xingshan ZengACM MM 2022 · 被引用 14 次
- CiCo: Domain-Aware Sign Language Retrieval via Cross-Lingual Contrastive LearningYiting Cheng, Fangyun Wei, Jianmin Bao, Dong Chen 等CVPR 2023
- A Simple Multi-Modality Transfer Learning Baseline for Sign Language TranslationYutong Chen, Fangyun Wei, Xiao Sun, Zhirong Wu 等CVPR 2022 · 被引用 137 次
- CVT-SLR: Contrastive Visual-Textual Transformation for Sign Language Recognition with Variational AlignmentJiangbin Zheng, Yile Wang, Cheng Tan, Siyuan Li 等CVPR 2023
- Graph Disentangled Contrastive Learning with Personalized Transfer for Cross-Domain RecommendationJing Liu, Lele Sun, Weizhi Nie, Peiguang Jing 等AAAI 2024 · 被引用 32 次
