SATO: Stable Text-to-Motion Framework
Wenshuo Chen, Hongru Xiao, Erhang Zhang, Lijie Hu, Lei Wang, Mengyuan Liu, Chen Chen
摘要
Is the Text to Motion model robust? Recent advancements in Text to Motion models primarily stem from more accurate predictions of specific actions. However, the text modality typically relies solely on pre-trained Contrastive Language-Image Pretraining (CLIP) models. Our research has uncovered a significant issue with the text-tomotion model: its predictions often exhibit inconsistent outputs, resulting in vastly different or even incorrect poses when presented with semantically similar or identical text inputs. In this paper, we undertake an analysis to elucidate the underlying causes of this instability, establishing a clear link between the unpredictability of model outputs and the erratic attention patterns of the text encoder module. Consequently, we introduce a formal framework aimed at addressing this issue, which we term the Stable Text-to-Motion Framework (SATO). SATO consists of three modules, each dedicated to stable attention, stable prediction, and maintaining a balance between accuracy and robustness trade-off. We present a methodology for constructing an SATO that satisfies the stability of attention and prediction. To verify the stability of the model, we introduced a new textual synonym perturbation dataset based on HumanML3D and KIT-ML. Results show that SATO is significantly more stable against synonyms and other slight perturbations while keeping its high accuracy performance. Codes and models are released at https://github.com/sato-team/Stable-Text-to-Motion-Framework
• Computing methodologies → Computer vision.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content GenerationKaishen Yuan, Yuting Zhang, Shang Gao, Yijie Zhu 等ICLR 2026 · 被引用 10 次
- Auto DragGAN: Editing the Generative Image Manifold in an Autoregressive MannerPengxiang Cai, Zhiwei Liu, Guibo Zhu, Yunfang Niu 等ACM MM 2024 · 被引用 3 次
- Towards Robust and Controllable Text-to-Motion via Masked Autoregressive DiffusionZongye Zhang, Bohan Kong, Qingjie Liu, Yunhong WangACM MM 2025 · 被引用 2 次
- ANT: Adaptive Neural Temporal-Aware Text-to-Motion ModelWenshuo Chen, Kuimou Yu, Haozhe Jia, Kaishen Yuan 等ACM MM 2025 · 被引用 1 次
- ChairPose: Pressure-based Chair Morphology Grounded Sitting Pose Estimation through Simulation-Assisted TrainingLala Shakti Swarup Ray, Vítor Fortes Rey, Bo Zhou, Paul Lukowicz 等UIST 2025 · 被引用 1 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 被引用 672 次
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang 等CVPR 2022 · 被引用 462 次
- ReMoDiffuse: Retrieval-Augmented Motion Diffusion ModelMingyuan Zhang, Xinying Guo, Liang Pan, Zhongang Cai 等ICCV 2023 · 被引用 301 次
相关 Paper
- Generative Motion Stylization of Cross-structure Characters within Canonical Motion SpaceJiaxu Zhang, Xin Chen, Gang Yu, Zhigang TuACM MM 2024 · 被引用 9 次
- TriC-Motion: Tri-Domain Causal Modeling Grounded Text-to-Motion GenerationYiyang Cao, Yunze Deng, Ziyu Lin, Bin Feng 等ICLR 2026
- Motion-Aligned Word Embeddings for Text-to-Motion GenerationKe Han, Yueming Lyu, Nicu SebeICLR 2026
- TMR: Text-to-Motion Retrieval Using Contrastive 3D Human Motion SynthesisMathis Petrovich, Michael J. Black, Gül VarolICCV 2023 · 被引用 192 次
- Reenact Anything: Semantic Video Motion Transfer Using Motion-Textual InversionManuel Kansy, Jacek Naruniec, Christopher Schroers, Markus Gross 等SIGGRAPH 2025 · 被引用 5 次
