Stable Signer: Hierarchical Sign Language Generative Model
Sen Fang, Yalin Feng, Hongbin Zhong, Yanxin Zhang, Dimitris N. Metaxas
Abstract
Sign Language Production (SLP) is the process of converting the complex input text into a real video. Most previous works focused on the Text2Gloss, Gloss2Pose, Pose2Vid stages 1 , and some concentrated on Prompt2Gloss and Text2Avatar stages. However, this field has made slow progress due to the inaccuracy of text conversion, pose generation, and the rendering of poses into real human videos in these stages, resulting in gradually accumulating errors. Therefore, in this paper, we streamline the traditional redundant structure, simplify and optimize the task objective, and design a new sign language generative model called Stable Signer. It redefines the SLP task as a hierarchical generation end-to-end task that only includes text understanding (Prompt2Gloss, Text2Gloss) and Pose2Vid, and executes text understanding through our proposed new Sign Language Understanding Linker called SLUL, and generates hand gestures through the named SLP-MoE hand gesture rendering expert block to end-to-end generate high-quality and multistyle sign language videos. SLUL is trained using the newly developed Semantic-Aware Gloss Masking Loss (SAGM Loss). Its performance has improved by 48.6% compared to the current SOTA generation methods, which is a significant increase in the SLP field. * Equal Contribution. 1 Gloss represents the specific gesture text words. The students will study in the library tomorrow. DWPose Denoising Off STUDENT LIBRARY STUDY TOMORROW Our clean pose input can be trained end-to-end with high quality. Random poses can cause noise that leads to blurriness. Introducing rule priors can reduce posture flickering. Minimizing the instability of the posture can reduce the burden of optimization. Denoising ON The semantic understanding usually involves Prompt2Gloss, Text2Gloss, and Gloss is generally in capitalized text.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df64a628-a9a6-4de8-8272-1d2ff9087a16Builds on16
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
- Follow Your Pose: Pose-Guided Text-to-Video Generation Using Pose-Free VideosYue Ma, Yingqing He, Xiaodong Cun, Xintao Wang et al.AAAI 2024 · 318 citations
- Human Motion Diffusion ModelGuy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir et al.ICLR 2023 · 167 citations
Related papers
- Gloss Semantic-Enhanced Network with Online Back-Translation for Sign Language ProductionShengeng Tang, Richang Hong, Dan Guo, Meng WangACM MM 2022 · 44 citations
- Towards Fast and High-Quality Sign Language ProductionWencan Huang, Wenwen Pan, Zhou Zhao, Qi TianACM MM 2021 · 43 citations
- SignPR: A Progressive Vector-Quantized Diffusion Framework for Sign Language ProductionXiao Liu, Shiwei Gan, Yafeng Yin, Bowen Guo et al.CVPR 2026 · 2 citations
- Focal-General Diffusion Model with Semantic Consistent Guidance for Sign Language ProductionYiheng Yu, Sheng Liu, Yuan Feng, Zhelun Jin et al.CVPR 2026
- DualSign: Semi-Supervised Sign Language Production with Balanced Multi-Modal Multi-Task Dual TransformationWencan Huang, Zhou Zhao, Jinzheng He, Mingmin ZhangACM MM 2022 · 8 citations
