Mixed SIGNals: Sign Language Production via a Mixture of Motion Primitives
Ben Saunders, Necati Cihan Camgöz, Richard Bowden
Abstract
It is common practice to represent spoken languages at their phonetic level. However, for sign languages, this implies breaking motion into its constituent motion primitives. Avatar based Sign Language Production (SLP) has traditionally done just this, building up animation from sequences of hand motions, shapes and facial expressions. However, more recent deep learning based solutions to SLP have tackled the problem using a single network that estimates the full skeletal structure. We propose splitting the SLP task into two distinct jointly-trained sub-tasks. The first translation sub-task translates from spoken language to a latent sign language representation, with gloss supervision. Subsequently, the animation sub-task aims to produce expressive sign language sequences that closely resemble the learnt spatiotemporal representation. Using a progressive transformer for the translation sub-task, we propose a novel Mixture of Motion Primitives (MOMP) architecture for sign language animation. A set of distinct motion primitives are learnt during training, that can be temporally combined at inference to animate continuous sign language sequences. We evaluate on the challenging RWTH-PHOENIX-Weather-2014T (PHOENIX14T) dataset, presenting extensive ablation studies and showing that MOMP outperforms baselines in user evaluations. We achieve state-of-the-art back translation performance with an 11% improvement over competing results. Importantly, and for the first time, we showcase stronger performance for a full translation pipeline going from spoken language to sign, than from gloss to sign.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 516279a7-aff5-4b5d-9ddd-4fce98874e55Cited by top-tier papers14
- Signing at Scale: Learning to Co-Articulate Signs for Large-Scale Photo-Realistic Sign Language ProductionBen Saunders, Necati Cihan Camgöz, Richard BowdenCVPR 2022 · 64 citations
- Gloss Semantic-Enhanced Network with Online Back-Translation for Sign Language ProductionShengeng Tang, Richang Hong, Dan Guo, Meng WangACM MM 2022 · 44 citations
- G2P-DDM: Generating Sign Pose Sequence from Gloss Sequence with Discrete Diffusion ModelPan Xie, Qipeng Zhang, Taiying Peng, Hao Tang et al.AAAI 2024 · 36 citations
- Advanced Sign Language Video Generation with Compressed and Quantized Multi-Condition TokenizationCong Wang, Zexuan Deng, Zhiwei Jiang, Yafeng Yin et al.NeurIPS 2025 · 13 citations
- Towards AI-driven Sign Language Generation with Non-manual MarkersHan Zhang, Rotem Shalev-Arkushin, Vasileios Baltatzis, Connor Gillis et al.CHI 2025 · 12 citations
Builds on4
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Towards Crowdsourced Training of Large Neural Networks using Decentralized Mixture-of-ExpertsMax Ryabinin, Anton GusevNeurIPS 2020 · 71 citations
- A Mixture of h - 1 Heads is Better than h HeadsHao Peng, Roy Schwartz, Dianqi Li, Noah A. SmithACL 2020 · 25 citations
- Sign Language Transformers: Joint End-to-End Sign Language Recognition and TranslationNecati Cihan Camgöz, Oscar Koller, Simon Hadfield, Richard BowdenCVPR 2020
Related papers
- DualSign: Semi-Supervised Sign Language Production with Balanced Multi-Modal Multi-Task Dual TransformationWencan Huang, Zhou Zhao, Jinzheng He, Mingmin ZhangACM MM 2022 · 8 citations
- Towards Fast and High-Quality Sign Language ProductionWencan Huang, Wenwen Pan, Zhou Zhao, Qi TianACM MM 2021 · 43 citations
- A Simple Multi-Modality Transfer Learning Baseline for Sign Language TranslationYutong Chen, Fangyun Wei, Xiao Sun, Zhirong Wu et al.CVPR 2022 · 137 citations
- Neural Sign Actors: A diffusion model for 3D sign language production from textVasileios Baltatzis, Rolandos Alexandros Potamias, Evangelos Ververas, Guanxiong Sun et al.CVPR 2024
- SignPR: A Progressive Vector-Quantized Diffusion Framework for Sign Language ProductionXiao Liu, Shiwei Gan, Yafeng Yin, Bowen Guo et al.CVPR 2026 · 2 citations
