Zero-Shot Text-to-Motion Evaluation using Video Language Models
Yuwen Ji, Donglin Wang, Yue Zhang
摘要
Text-to-motion (T2M) generation has become a fundamental task, yet existing evaluation metrics often fail to capture whether a generated motion semantically matches its text description. We propose VeMo, a zero-shot evaluation framework that renders generated human motions into videos and uses pretrained video-language models (VLMs) to assess text-motion alignment. Instead of training an evaluator on scarce motion-specific labels, VeMo transfers the semantic reasoning ability of VLMs to T2M evaluation through normalized likelihood-based scoring. To reduce the effect of 3D-to-2D projection ambiguity, we introduce an entropy-driven uncertainty analysis for identifying reliable rendered views. To address the lack of rigorous standards in the field, we further contribute a test-only and human-annotated meta-evaluation benchmark, covering motions generated by multiple representative T2M models. Extensive experiments show that VeMo correlates better with human judgments than existing reference-based and reference-free metrics. Additional analyses on view selection, rendering protocols, textual prompt robustness, and computational trade-offs characterize both the promise and limitations of VLM-based T2M evaluation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll 等ICCV 2019 · 被引用 1,784 次
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu 等NeurIPS 2023 · 被引用 698 次
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng 等EMNLP 2024 · 被引用 479 次
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li 等ICLR 2024 · 被引用 467 次
相关 Paper
- VMBench: A Benchmark for Perception-Aligned Video Motion GenerationXinran Ling, Chen Zhu, Meiqi Wu, Hangyu Li 等ICCV 2025 · 被引用 2 次
- ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language ModelsIlker Kesen, Andrea Pedrotti, Mustafa Dogan, Michele Cafagna 等ICLR 2024 · 被引用 25 次
- LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive TokensZekun Li, Sizhe An, Chengcheng Tang, Chuan Guo 等CVPR 2026 · 被引用 12 次
- Motion-Aligned Word Embeddings for Text-to-Motion GenerationKe Han, Yueming Lyu, Nicu SebeICLR 2026
- ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion TransferJiayi Gao, Zijin Yin, Changcheng Hua, Yuxin Peng 等CVPR 2025
