AniMo: Species-Aware Model for Text-Driven Animal Motion Generation
Xuan Wang, Kai Ruan, Xing Zhang, Gaoang Wang
Abstract
Text-driven motion generation has made significant strides in recent years. However, most existing works focus on human motion, largely overlooking the rich and diverse behaviors of animals. Understanding and synthesizing animal motion have important applications in wildlife conservation, animal ecology, and biomechanics. Animal motion modeling presents unique challenges due to species diversity, varied morphological structures, and different behavioral patterns in response to similar textual descriptions. To address these challenges, we propose AniMo for text-driven animal motion generation. AniMo consists of two stages: motion tokenization and text-to-motion generation. In the motion tokenization stage, we encode motions using a jointaware spatiotemporal encoder with species-aware feature modulation, enabling the model to adapt to diverse skeletal structures across species. In the text-to-motion generation stage, we employ masked modeling to jointly learn the mapping from textual descriptions to motion tokens. Additionally, we introduce AniMo4D, a large-scale dataset containing 78,149 motion sequences and 185,435 textual descriptions across 114 animal species. Experimental results show that AniMo achieves superior performance on both the AniMo4D and AnimalML3D datasets, effectively capturing diverse morphological structures and behavioral patterns across animal species.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b6665fe-9137-4fdf-998f-612913465da4Cited by top-tier papers3
- Semantic-Aware Motion Encoding for Topology-Agnostic Character AnimationZongye Zhang, Yuzhuo Cui, Qingjie Liu, Yunhong WangICML 2026 · 1 citation
- FunPhase: A Periodic Functional Autoencoder for Motion Generation via Phase ManifoldsMarco Pegoraro, Evan Atherton, Bruno Roy, Aliasghar Khani et al.ICML 2026 · 1 citation
- X-MoGen: Unified Motion Generation Across Humans and AnimalsXuan Wang, Kai Ruan, Liyang Qian, Guo Zhi Zhi et al.AAAI 2026
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar et al.NeurIPS 2023 · 910 citations
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu et al.NeurIPS 2023 · 698 citations
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang et al.CVPR 2022 · 462 citations
- Finite Scalar Quantization: VQ-VAE Made SimpleFabian Mentzer, David Minnen, Eirikur Agustsson, Michael TschannenICLR 2024 · 442 citations
Related papers
- OmniMotionGPT: Animal Motion Generation with Limited DataZhangsihao Yang, Mingyuan Zhou, Mengyi Shan, Bingbing Wen et al.CVPR 2024 · 6 citations
- SnapMoGen: Human Motion Generation from Expressive TextsChuan Guo, Inwoo Hwang, Jian Wang, Bing ZhouNeurIPS 2025 · 50 citations
- GENMO: A GENeralist Model for Human MOtionJiefeng Li, Jinkun Cao, Haotian Zhang, Davis Rempe et al.ICCV 2025 · 15 citations
- How to Move Your Dragon: Text-to-Motion Synthesis for Large-Vocabulary ObjectsWonkwang Lee, Jongwon Jeong, Taehong Moon, Hyeon-Jong Kim et al.ICML 2025
- RigMo: Unifying Rig and Motion Learning for Generative AnimationHao Zhang, Jiahao Luo, Bohui Wan, Yizhou Zhao et al.CVPR 2026 · 6 citations
