Greg: GEometry-Aware RegIon Refinement for Sign Language Video Generation
Tongkai Shi, Lianyu Hu, Fanhua Shang, Liqing Gao, Wei Feng
Abstract
Sign Language Video Generation (SLVG) aims to transform sign language sequences into natural and fluent sign language videos. Existing SLVG methods lack geometric modeling of human anatomical structures, leading to anatomically implausible and temporally inconsistent generation. To address these challenges, we propose a novel framework: Geometry-Aware Region Refinement (GReg) for SLVG. GReg uses geometric information (such as normal maps and gradient maps) from the SMPL-X model to ensure anatomical and temporal consistency. To fully leverage the geometric priors, we propose two novel methods: 1) Regional Prior Generation, which uses regional expert networks to generate target-structured regions as generation priors; 2) Gradient-enhanced Refinement, which guides the refinement of detailed structures in key regions using gradient features. Furthermore, we enhance visual realism in key regions through adversarial training on both these regions and their gradient maps. Experimental results demonstrate that GReg achieves state-of-the-art performance with superior structural accuracy and temporal consistency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on13
- Two-Stream Network for Sign Language Recognition and TranslationYutong Chen, Ronglai Zuo, Fangyun Wei, Yu Wu et al.NeurIPS 2022 · 288 citations
- Thin-Plate Spline Motion Model for Image AnimationJian Zhao, Hui ZhangCVPR 2022 · 196 citations
- Mixed SIGNals: Sign Language Production via a Mixture of Motion PrimitivesBen Saunders, Necati Cihan Camgöz, Richard BowdenICCV 2021 · 82 citations
- Neural Texture Extraction and Distribution for Controllable Person Image SynthesisYurui Ren, Xiaoqing Fan, Ge Li, Shan Liu et al.CVPR 2022 · 81 citations
- Signing at Scale: Learning to Co-Articulate Signs for Large-Scale Photo-Realistic Sign Language ProductionBen Saunders, Necati Cihan Camgöz, Richard BowdenCVPR 2022 · 64 citations
Related papers
- Advanced Sign Language Video Generation with Compressed and Quantized Multi-Condition TokenizationCong Wang, Zexuan Deng, Zhiwei Jiang, Yafeng Yin et al.NeurIPS 2025 · 13 citations
- SignPR: A Progressive Vector-Quantized Diffusion Framework for Sign Language ProductionXiao Liu, Shiwei Gan, Yafeng Yin, Bowen Guo et al.CVPR 2026 · 2 citations
- Discrete to Continuous: Generating Smooth Transition Poses from Sign Language ObservationsShengeng Tang, Jiayi He, Lechao Cheng, Jingjing Wu et al.CVPR 2025
- TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language TranslationDongxu Li, Chenchen Xu, Xin Yu, Kaihao Zhang et al.NeurIPS 2020 · 171 citations
- Animatable Virtual Humans: Learning Pose-Dependent Human Representations in UV Space for Interactive Performance SynthesisWieland Morgenstern, Milena T. Bagdasarian, Anna Hilsmann, Peter EisertIEEE VR 2024 · 7 citations
