AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision Reward
Haonan Han, Xiangzuo Wu, Huan Liao, Zunnan Xu, Zhongyuan Hu, Ronghui Li, Yachao Zhang, Xiu Li
Abstract
A person walks forward, bends at the waist, and picks up something" "A person waves his arm first, walks forward in a straight line, then turns left." "A person leaps forward for 3 times" Left Right Left Right Left Right : Left is better : Left is better : Right is better Integrity Temporal Frequency Figure 1. Showcases of motion samples for three scenarios. The two motion samples for each scenario were generated based on the prompt above the samples. Moreover, we leverage GPT-4V to compare two motion samples according to the degree of alignment between the motion samples and the input prompt.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02065e2f-ffdd-4323-b415-16ae04e4676eCited by top-tier papers7
- LottieGPT: Tokenizing Vector Animation for Autoregressive GenerationJunhao Chen, Kejun Gao, Yuehan Cui, Mingze Sun et al.CVPR 2026 · 10 citations
- Separate to Collaborate: Dual-Stream Diffusion Model for Coordinated Piano Hand Motion SynthesisZihao Liu, Mingwen Ou, Zunnan Xu, Jiaqi Huang et al.ACM MM 2025 · 2 citations
- U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generationxiang deng, Feng Gao, Yong Zhang, Youxin Pang et al.CVPR 2026 · 2 citations
- GazeInterpreter: Parsing Eye Gaze to Generate Eye-Body-Coordinated NarrationsQing Chang, Zhiming HuAAAI 2026
- Zero-Shot Text-to-Motion Evaluation using Video Language ModelsYuwen Ji, Donglin Wang, Yue ZhangICML 2026
Builds on12
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong et al.NeurIPS 2023 · 1,310 citations
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu et al.NeurIPS 2023 · 698 citations
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang et al.CVPR 2022 · 462 citations
- Human Motion Diffusion as a Generative PriorYoni Shafir, Guy Tevet, Roy Kapon, Amit Haim BermanoICLR 2024 · 371 citations
Related papers
- MoMask: Generative Masked Modeling of 3D Human MotionsChuan Guo, Yuxuan Mu, Muhammad Gohar Javed, Sen Wang et al.CVPR 2024
- SINC: Spatial Composition of 3D Human Motions for Simultaneous Action GenerationNikos Athanasiou, Mathis Petrovich, Michael J. Black, Gül VarolICCV 2023 · 69 citations
- VMBench: A Benchmark for Perception-Aligned Video Motion GenerationXinran Ling, Chen Zhu, Meiqi Wu, Hangyu Li et al.ICCV 2025 · 2 citations
- HSI-GPT: A General-Purpose Large Scene-Motion-Language Model for Human Scene InteractionYuan Wang, Yali Li, Xiang Li, Shengjin WangCVPR 2025
- ReMoGPT: Part-Level Retrieval-Augmented Motion-Language ModelsQing Yu, Mikihiro Tanaka, Kent FujiwaraAAAI 2025 · 6 citations
