SEEG: Semantic Energized Co-speech Gesture Generation
Yuanzhi Liang, Qianyu Feng, Linchao Zhu, Li Hu, Pan Pan, Yi Yang
摘要
Talking gesture generation is a practical yet challenging task that aims to synthesize gestures in line with speech. Gestures with meaningful signs can better convey useful information and arouse sympathy in the audience. Current works focus on aligning gestures with the speech rhythms, which are difficult to mine the semantics and model semantic gestures explicitly. This paper proposes a novel SEmantic Energized Generation (SEEG) method for semanticaware gesture generation. Our method contains two parts: DEcoupled Mining module (DEM) and Semantic Energizing Module (SEM). DEM decouples the semantic-irrelevant information from inputs and separately mines information for the beat and semantic gestures. SEM conducts semantic learning and produces semantic gestures. Apart from representational similarity, SEM requires the predictions to express the same semantics as the ground truth. Besides, a semantic prompter is designed in SEM to leverage the semantic-aware supervision to predictions. This promotes the networks to learn and generate semantic gestures. Experimental results reported in three metrics on different benchmarks prove that SEEG efficiently mines semantic cues and generates semantic gestures. SEEG outperforms other methods in all semantic-aware evaluations on different datasets. Qualitative evaluations also indicate the superiority of SEEG in semantic expressiveness. Code is available via https://github.com/akira-l/SEEG . * This work was performed at Alibaba DAMO Academy, Alibaba Group. (a) Diverse and expressive semantic gestures (b) Intuitive and semantic-irrelevant beat gestures To launch• • • • • • • a big • • • • • • • • • investigation Know • • that • • we • • • • • can • • • • • make • • • • • a decision
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- GestureDiffuCLIP: Gesture Diffusion Model with CLIP LatentsTenglong Ao, Zeyi Zhang, Libin LiuSIGGRAPH 2023 · 被引用 151 次
- LivelySpeaker: Towards Semantic-Aware Co-Speech Gesture GenerationYihao Zhi, Xiaodong Cun, Xuelin Chen, Xi Shen 等ICCV 2023 · 被引用 48 次
- Semantic Gesticulator: Semantics-Aware Co-Speech Gesture SynthesisZeyi Zhang, Tenglong Ao, Yuyao Zhang, Qingzhe Gao 等SIGGRAPH 2024 · 被引用 39 次
- Modality-Independent Teachers Meet Weakly-Supervised Audio-Visual Event ParserYung-Hsuan Lai, Yen-Chun Chen, Frank WangNeurIPS 2023 · 被引用 27 次
- UnifiedGesture: A Unified Gesture Synthesis Model for Multiple SkeletonsSicheng Yang, Zilin Wang, Zhiyong Wu, Minglei Li 等ACM MM 2023 · 被引用 17 次
它引用的顶会 Paper3
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Speech Drives Templates: Co-Speech Gesture Synthesis with Learned TemplatesShenhan Qian, Zhi Tu, Yihao Zhi, Wen Liu 等ICCV 2021 · 被引用 95 次
- Body2Hands: Learning To Infer 3D Hands From Conversational Gesture Body DynamicsEvonne Ng, Shiry Ginosar, Trevor Darrell, Hanbyul JooCVPR 2021
相关 Paper
- SemTalk: Holistic Co-Speech Motion Generation with Frame-Level Semantic EmphasisXiangyue Zhang, Jianfang Li, Jiaxu Zhang, Ziqiang Dang 等ICCV 2025 · 被引用 12 次
- SemGesture: Synthesizing Semantically Enhanced and Coherent GesturesPengsheng Liu, Zhaojie Chu, Xiaofen Xing, Xiangmin XuACM MM 2025 · 被引用 2 次
- SemGes: Semantics-Aware Co-Speech Gesture Generation Using Semantic Coherence and Relevance LearningLanmiao Liu, Esam Ghaleb, Asli Özyürek, Zerrin YumakICCV 2025 · 被引用 4 次
- Retrieving Semantics from the Deep: an RAG Solution for Gesture SynthesisMuhammad Hamza Mughal, Rishabh Dabral, Merel C. J. Scholman, Vera Demberg 等CVPR 2025
- Emphasizing Semantic Consistency of Salient Posture for Speech-Driven Gesture GenerationFengqi Liu, Hexiang Wang, Jingyu Gong, Ran Yi 等ACM MM 2024 · 被引用 2 次
