Lune

CVPR2024顶会

AvatarGPT: All-in-One Framework for Motion Understanding, Planning, Generation and Beyond

Zixiang Zhou, Yu Wan, Baoyuan Wang

2024年份
19被引次数
31顶会引用

摘要

Large Language Models(LLMs) have shown remarkable emergent abilities in unifying almost all (if not every) NLP tasks. In the human motion-related realm, however, researchers still develop siloed models for each task. In-spired by InstuctGPT[16], and the generalist concept be-hind Gato [27], we introduce AvatarGPT, an All-in-One framework for motion understanding, planning, generations as well as other tasks such as motion in-between synthesis. AvatarGPT treats each task as one type of in-struction fine-tuned on the shared LLM. All the tasks are seamlessly interconnected with language as the univer-sal interface, constituting a closed-loop within the frame-work. To achieve this, human motion sequences are first encoded as discrete tokens, which serve as the extended vo-cabulary of LLM. Then, an unsupervised pipeline to gen-erate natural language descriptions of human action sequences from in-the-wild videos is developed. Finally, all tasks are jointly trained. Extensive experiments show that AvatarGPT achieves SOTA on low-level tasks, and promising results on high-level tasks, demonstrating the effectiveness of our proposed All-in-One framework. Moreover, for the first time, AvatarGPT enables a principled approach by iterative traversal of the tasks within the closed-loop for un-limited long-motion synthesis.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper31

问问它们各自怎么用它

它引用的顶会 Paper23

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖