SATO: Stable Text-to-Motion Framework
Wenshuo Chen, Hongru Xiao, Erhang Zhang, Lijie Hu, Lei Wang, Mengyuan Liu, Chen Chen
Abstract
Is the Text to Motion model robust? Recent advancements in Text to Motion models primarily stem from more accurate predictions of specific actions. However, the text modality typically relies solely on pre-trained Contrastive Language-Image Pretraining (CLIP) models. Our research has uncovered a significant issue with the text-tomotion model: its predictions often exhibit inconsistent outputs, resulting in vastly different or even incorrect poses when presented with semantically similar or identical text inputs. In this paper, we undertake an analysis to elucidate the underlying causes of this instability, establishing a clear link between the unpredictability of model outputs and the erratic attention patterns of the text encoder module. Consequently, we introduce a formal framework aimed at addressing this issue, which we term the Stable Text-to-Motion Framework (SATO). SATO consists of three modules, each dedicated to stable attention, stable prediction, and maintaining a balance between accuracy and robustness trade-off. We present a methodology for constructing an SATO that satisfies the stability of attention and prediction. To verify the stability of the model, we introduced a new textual synonym perturbation dataset based on HumanML3D and KIT-ML. Results show that SATO is significantly more stable against synonyms and other slight perturbations while keeping its high accuracy performance. Codes and models are released at https://github.com/sato-team/Stable-Text-to-Motion-Framework
• Computing methodologies → Computer vision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47330af4-bce4-4f2e-832e-a13012374e98Cited by top-tier papers5
- CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content GenerationKaishen Yuan, Yuting Zhang, Shang Gao, Yijie Zhu et al.ICLR 2026 · 10 citations
- Auto DragGAN: Editing the Generative Image Manifold in an Autoregressive MannerPengxiang Cai, Zhiwei Liu, Guibo Zhu, Yunfang Niu et al.ACM MM 2024 · 3 citations
- Towards Robust and Controllable Text-to-Motion via Masked Autoregressive DiffusionZongye Zhang, Bohan Kong, Qingjie Liu, Yunhong WangACM MM 2025 · 2 citations
- ANT: Adaptive Neural Temporal-Aware Text-to-Motion ModelWenshuo Chen, Kuimou Yu, Haozhe Jia, Kaishen Yuan et al.ACM MM 2025 · 1 citation
- ChairPose: Pressure-based Chair Morphology Grounded Sitting Pose Estimation through Simulation-Assisted TrainingLala Shakti Swarup Ray, Vítor Fortes Rey, Bo Zhou, Paul Lukowicz et al.UIST 2025 · 1 citation
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 672 citations
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang et al.CVPR 2022 · 462 citations
- ReMoDiffuse: Retrieval-Augmented Motion Diffusion ModelMingyuan Zhang, Xinying Guo, Liang Pan, Zhongang Cai et al.ICCV 2023 · 301 citations
Related papers
- Generative Motion Stylization of Cross-structure Characters within Canonical Motion SpaceJiaxu Zhang, Xin Chen, Gang Yu, Zhigang TuACM MM 2024 · 9 citations
- TriC-Motion: Tri-Domain Causal Modeling Grounded Text-to-Motion GenerationYiyang Cao, Yunze Deng, Ziyu Lin, Bin Feng et al.ICLR 2026
- Motion-Aligned Word Embeddings for Text-to-Motion GenerationKe Han, Yueming Lyu, Nicu SebeICLR 2026
- TMR: Text-to-Motion Retrieval Using Contrastive 3D Human Motion SynthesisMathis Petrovich, Michael J. Black, Gül VarolICCV 2023 · 192 citations
- Reenact Anything: Semantic Video Motion Transfer Using Motion-Textual InversionManuel Kansy, Jacek Naruniec, Christopher Schroers, Markus Gross et al.SIGGRAPH 2025 · 5 citations
