Greg: GEometry-Aware RegIon Refinement for Sign Language Video Generation
Tongkai Shi, Lianyu Hu, Fanhua Shang, Liqing Gao, Wei Feng
摘要
Sign Language Video Generation (SLVG) aims to transform sign language sequences into natural and fluent sign language videos. Existing SLVG methods lack geometric modeling of human anatomical structures, leading to anatomically implausible and temporally inconsistent generation. To address these challenges, we propose a novel framework: Geometry-Aware Region Refinement (GReg) for SLVG. GReg uses geometric information (such as normal maps and gradient maps) from the SMPL-X model to ensure anatomical and temporal consistency. To fully leverage the geometric priors, we propose two novel methods: 1) Regional Prior Generation, which uses regional expert networks to generate target-structured regions as generation priors; 2) Gradient-enhanced Refinement, which guides the refinement of detailed structures in key regions using gradient features. Furthermore, we enhance visual realism in key regions through adversarial training on both these regions and their gradient maps. Experimental results demonstrate that GReg achieves state-of-the-art performance with superior structural accuracy and temporal consistency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Two-Stream Network for Sign Language Recognition and TranslationYutong Chen, Ronglai Zuo, Fangyun Wei, Yu Wu 等NeurIPS 2022 · 被引用 288 次
- Thin-Plate Spline Motion Model for Image AnimationJian Zhao, Hui ZhangCVPR 2022 · 被引用 196 次
- Mixed SIGNals: Sign Language Production via a Mixture of Motion PrimitivesBen Saunders, Necati Cihan Camgöz, Richard BowdenICCV 2021 · 被引用 82 次
- Neural Texture Extraction and Distribution for Controllable Person Image SynthesisYurui Ren, Xiaoqing Fan, Ge Li, Shan Liu 等CVPR 2022 · 被引用 81 次
- Signing at Scale: Learning to Co-Articulate Signs for Large-Scale Photo-Realistic Sign Language ProductionBen Saunders, Necati Cihan Camgöz, Richard BowdenCVPR 2022 · 被引用 64 次
相关 Paper
- Advanced Sign Language Video Generation with Compressed and Quantized Multi-Condition TokenizationCong Wang, Zexuan Deng, Zhiwei Jiang, Yafeng Yin 等NeurIPS 2025 · 被引用 13 次
- SignPR: A Progressive Vector-Quantized Diffusion Framework for Sign Language ProductionXiao Liu, Shiwei Gan, Yafeng Yin, Bowen Guo 等CVPR 2026 · 被引用 2 次
- Discrete to Continuous: Generating Smooth Transition Poses from Sign Language ObservationsShengeng Tang, Jiayi He, Lechao Cheng, Jingjing Wu 等CVPR 2025
- TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language TranslationDongxu Li, Chenchen Xu, Xin Yu, Kaihao Zhang 等NeurIPS 2020 · 被引用 171 次
- Animatable Virtual Humans: Learning Pose-Dependent Human Representations in UV Space for Interactive Performance SynthesisWieland Morgenstern, Milena T. Bagdasarian, Anna Hilsmann, Peter EisertIEEE VR 2024 · 被引用 7 次
