Towards Fast and High-Quality Sign Language Production
Wencan Huang, Wenwen Pan, Zhou Zhao, Qi Tian
Abstract
Sign Language Production (SLP) aims to automatically translate a spoken language description to its corresponding sign language video. The core procedure of SLP is to transform sign gloss intermediaries into sign pose sequences (G2P). Most existing methods for G2P are based on sequential autoregression or sequence-to-sequence encoder-decoder learning. However, by generating target pose frames conditioned on the previously generated ones, these models are prone to bringing issues such as error accumulation and high inference latency. In this paper, we argue that such issues are mainly caused by adopting autoregressive manner. Hence, we propose a novel Non-AuToregressive (NAT) model with a parallel decoding scheme, as well as an External Aligner for sequence alignment learning. Specifically, we extract alignments from the external aligner by monotonic alignment search for gloss duration prediction, which is used by a length regulator to expand the source gloss sequence to match the length of the target sign pose sequence for parallel sign pose generation. Furthermore, we devise a spatial-temporal graph convolutional pose generator in the NAT model to generate smoother and more natural sign pose sequences. Extensive experiments conducted on PHOENIX14T dataset show that our proposed model outperforms state-of-the-art autoregressive models in terms of speed and quality.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b1b864c0-b828-481d-a697-c14ff3a3dcb3Cited by top-tier papers12
- Signing at Scale: Learning to Co-Articulate Signs for Large-Scale Photo-Realistic Sign Language ProductionBen Saunders, Necati Cihan Camgöz, Richard BowdenCVPR 2022 · 64 citations
- Gloss Semantic-Enhanced Network with Online Back-Translation for Sign Language ProductionShengeng Tang, Richang Hong, Dan Guo, Meng WangACM MM 2022 · 44 citations
- SimulSLT: End-to-End Simultaneous Sign Language TranslationAoxiong Yin, Zhou Zhao, Jinglin Liu, Weike Jin et al.ACM MM 2021 · 35 citations
- Sign-IDD: Iconicity Disentangled Diffusion for Sign Language ProductionShengeng Tang, Jiayi He, Dan Guo, Yanyan Wei et al.AAAI 2025 · 23 citations
- Signs as Tokens: A Retrieval-Enhanced Multilingual Sign Language GeneratorRonglai Zuo, Rolandos Alexandros Potamias, Evangelos Ververas, Jiankang Deng et al.ICCV 2025 · 9 citations
Related papers
- G2P-DDM: Generating Sign Pose Sequence from Gloss Sequence with Discrete Diffusion ModelPan Xie, Qipeng Zhang, Taiying Peng, Hao Tang et al.AAAI 2024 · 36 citations
- Focal-General Diffusion Model with Semantic Consistent Guidance for Sign Language ProductionYiheng Yu, Sheng Liu, Yuan Feng, Zhelun Jin et al.CVPR 2026
- SignPR: A Progressive Vector-Quantized Diffusion Framework for Sign Language ProductionXiao Liu, Shiwei Gan, Yafeng Yin, Bowen Guo et al.CVPR 2026 · 2 citations
- C2ST: Cross-modal Contextualized Sequence Transduction for Continuous Sign Language RecognitionHuaiwen Zhang, Zihang Guo, Yang Yang, Xin Liu et al.ICCV 2023 · 21 citations
- Mixed SIGNals: Sign Language Production via a Mixture of Motion PrimitivesBen Saunders, Necati Cihan Camgöz, Richard BowdenICCV 2021 · 82 citations
