SignPR: A Progressive Vector-Quantized Diffusion Framework for Sign Language Production
Xiao Liu, Shiwei Gan, Yafeng Yin, Bowen Guo, Zhiwei Jiang, Shunmei Meng, Lei Xie, Sanglu Lu
Abstract
Sign language production aims to generate sign sequences from spoken language, where the generation of sign pose sequences from text is often treated as a significant task. However, due to the differences in grammatical rules and modalities between sign language pose sequences and spoken language text, it is rather challenging to convert text into sign poses (i.e., Text2Pose), while maintaining semantic consistency, motion accuracy and temporal coherence. In this paper, we focus on the Text2Pose task, and propose SignPR, a progressive diffusion framework that jointly models the structural and temporal properties of signing. Structurally, we perform progressive structural refinement: a structural VQVAE encodes each frame into semantic-aware and region-based discrete representations; the diffusion process first produces semantically consistent poses, and then progressively refines motion details under text and semantic conditioning. Temporally, we introduce block-wise causal diffusion, which progressively enforces temporal coherence and enables iterative refinement to earlier generated segments, yielding smoother transitions and reduced jitter. Extensive experiments on widely used datasets demonstrate that SignPR achieves superior results compared with prior T2P methods across multiple metrics, producing pose sequences that are semantically faithful, motion-accurate, and temporally coherent.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d1ac559-6154-4832-875d-02fccfa08cf1Builds on21
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Vector Quantized Diffusion Model for Text-to-Image SynthesisShuyang Gu, Dong Chen, Jianmin Bao, Fang Wen et al.CVPR 2022 · 607 citations
Related papers
- G2P-DDM: Generating Sign Pose Sequence from Gloss Sequence with Discrete Diffusion ModelPan Xie, Qipeng Zhang, Taiying Peng, Hao Tang et al.AAAI 2024 · 36 citations
- Focal-General Diffusion Model with Semantic Consistent Guidance for Sign Language ProductionYiheng Yu, Sheng Liu, Yuan Feng, Zhelun Jin et al.CVPR 2026
- Sign-IDD: Iconicity Disentangled Diffusion for Sign Language ProductionShengeng Tang, Jiayi He, Dan Guo, Yanyan Wei et al.AAAI 2025 · 23 citations
- Towards Fast and High-Quality Sign Language ProductionWencan Huang, Wenwen Pan, Zhou Zhao, Qi TianACM MM 2021 · 43 citations
- Discrete to Continuous: Generating Smooth Transition Poses from Sign Language ObservationsShengeng Tang, Jiayi He, Lechao Cheng, Jingjing Wu et al.CVPR 2025
