Sign-IDD: Iconicity Disentangled Diffusion for Sign Language Production
Shengeng Tang, Jiayi He, Dan Guo, Yanyan Wei, Feng Li, Richang Hong
Abstract
Sign Language Production (SLP) aims to generate semantically consistent sign videos from textual statements, where the conversion from textual glosses to sign poses (G2P) is a crucial step. Existing G2P methods typically treat sign poses as discrete three-dimensional coordinates and directly fit them, which overlooks the relative positional relationships among joints. To this end, we provide a new perspective, constraining joint associations and gesture details by modeling the limb bones to improve the accuracy and naturalness of the generated poses. In this work, we propose a pioneering iconicity disentangled diffusion framework, termed Sign-IDD, specifically designed for SLP. Sign-IDD incorporates a novel Iconicity Disentanglement (ID) module to bridge the gap between relative positions among joints. The ID module disentangles the conventional 3D joint representation into a 4D bone representation, comprising the 3D spatial direction vector and 1D spatial distance vector between adjacent joints. Additionally, an Attribute Controllable Diffusion (ACD) module is introduced to further constrain joint associations, in which the attribute separation layer aims to separate the bone direction and length attributes, and the attribute control layer is designed to guide the pose generation by leveraging the above attributes. The ACD module utilizes the gloss embeddings as semantic conditions and finally generates sign poses from noise embeddings. Extensive experiments on PHOENIX14T and USTC-CSL datasets validate the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f362b5f-757b-49ed-ab7a-043b0405c67eCited by top-tier papers5
- Signs as Tokens: A Retrieval-Enhanced Multilingual Sign Language GeneratorRonglai Zuo, Rolandos Alexandros Potamias, Evangelos Ververas, Jiankang Deng et al.ICCV 2025 · 9 citations
- SignPR: A Progressive Vector-Quantized Diffusion Framework for Sign Language ProductionXiao Liu, Shiwei Gan, Yafeng Yin, Bowen Guo et al.CVPR 2026 · 2 citations
- Wi-CBR: Salient-aware Adaptive WiFi Sensing for Cross-domain Behavior RecognitionRuobei Zhang, Shengeng Tang, Huan Yan, Xiang Zhang et al.AAAI 2026 · 2 citations
- Focal-General Diffusion Model with Semantic Consistent Guidance for Sign Language ProductionYiheng Yu, Sheng Liu, Yuan Feng, Zhelun Jin et al.CVPR 2026
- OmniVL-Guard: Towards Unified Vision-Language Forgery Detection and Grounding via Balanced RLJinjie Shen, Jing Wu, Yaxiong Wang, Lechao Cheng et al.ICML 2026
Builds on12
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang et al.NeurIPS 2022 · 1,546 citations
- ResDiff: Combining CNN and Diffusion Model for Image Super-resolutionShuyao Shang, Zhengyang Shan, Guangxing Liu, Lunqian Wang et al.AAAI 2024 · 158 citations
- Diffusion-Based 3D Human Pose Estimation with Multi-Hypothesis AggregationWenkang Shan, Zhenhua Liu, Xinfeng Zhang, Zhao Wang et al.ICCV 2023 · 148 citations
- Signing at Scale: Learning to Co-Articulate Signs for Large-Scale Photo-Realistic Sign Language ProductionBen Saunders, Necati Cihan Camgöz, Richard BowdenCVPR 2022 · 64 citations
Related papers
- Towards Fast and High-Quality Sign Language ProductionWencan Huang, Wenwen Pan, Zhou Zhao, Qi TianACM MM 2021 · 43 citations
- G2P-DDM: Generating Sign Pose Sequence from Gloss Sequence with Discrete Diffusion ModelPan Xie, Qipeng Zhang, Taiying Peng, Hao Tang et al.AAAI 2024 · 36 citations
- Gloss Semantic-Enhanced Network with Online Back-Translation for Sign Language ProductionShengeng Tang, Richang Hong, Dan Guo, Meng WangACM MM 2022 · 44 citations
- Advanced Sign Language Video Generation with Compressed and Quantized Multi-Condition TokenizationCong Wang, Zexuan Deng, Zhiwei Jiang, Yafeng Yin et al.NeurIPS 2025 · 13 citations
- Discrete to Continuous: Generating Smooth Transition Poses from Sign Language ObservationsShengeng Tang, Jiayi He, Lechao Cheng, Jingjing Wu et al.CVPR 2025
