DualSign: Semi-Supervised Sign Language Production with Balanced Multi-Modal Multi-Task Dual Transformation
Wencan Huang, Zhou Zhao, Jinzheng He, Mingmin Zhang
摘要
Sign Language Production (SLP) aims to translate a spoken language description to its corresponding continuous sign language sequence. A prevailing solution for this problem is in a two-staged manner: it formulates SLP as two sub-tasks, i.e., Text to Gloss (T2G) translation and Gloss to Pose (G2P) animation, with gloss annotations as pivots. Although two-staged approaches achieve better performance than their direct translation counterparts, the requirement of gloss intermediaries causes a parallel data bottleneck. In this paper, to reduce reliance on gloss annotations in two-staged approaches, we propose DualSign, a semi-supervised two-staged SLP framework, which can effectively utilize partially gloss-annotated text-pose pairs and monolingual gloss data. The key component of DualSign is a novel Balanced Multi-Modal Multi-Task Dual Transformation (BM3T-DT) method, where two well-designed models, i.e., a Multi-Modal T2G model (MM-T2G) and a Multi-Task G2P model (MT-G2P), are jointly trained by leveraging their task duality and unlabeled data. After applying BM3T-DT, we derive the expected uni-modal T2G model from the well-trained MM-T2G with knowledge distillation. Considering that the MM-T2G may suffer from modality imbalance when decoding with multiple input modalities, we devise a cross-modal balancing loss, further boosting the system's overall performance. Extensive experiments conducted on the PHOENIX14T dataset show the effectiveness of our approach in the semi-supervised setting. By training with additionally collected unlabeled data, DualSign substantially improves previous state-of-the-art SLP methods.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Signs as Tokens: A Retrieval-Enhanced Multilingual Sign Language GeneratorRonglai Zuo, Rolandos Alexandros Potamias, Evangelos Ververas, Jiankang Deng 等ICCV 2025 · 被引用 9 次
- SignPR: A Progressive Vector-Quantized Diffusion Framework for Sign Language ProductionXiao Liu, Shiwei Gan, Yafeng Yin, Bowen Guo 等CVPR 2026 · 被引用 2 次
相关 Paper
- Gloss Semantic-Enhanced Network with Online Back-Translation for Sign Language ProductionShengeng Tang, Richang Hong, Dan Guo, Meng WangACM MM 2022 · 被引用 44 次
- Mixed SIGNals: Sign Language Production via a Mixture of Motion PrimitivesBen Saunders, Necati Cihan Camgöz, Richard BowdenICCV 2021 · 被引用 82 次
- Focal-General Diffusion Model with Semantic Consistent Guidance for Sign Language ProductionYiheng Yu, Sheng Liu, Yuan Feng, Zhelun Jin 等CVPR 2026
- A Simple Multi-Modality Transfer Learning Baseline for Sign Language TranslationYutong Chen, Fangyun Wei, Xiao Sun, Zhirong Wu 等CVPR 2022 · 被引用 137 次
- Towards Fast and High-Quality Sign Language ProductionWencan Huang, Wenwen Pan, Zhou Zhao, Qi TianACM MM 2021 · 被引用 43 次
