Humanoid Locomotion as Next Token Prediction
Ilija Radosavovic, Bike Zhang, Baifeng Shi, Jathushan Rajasegaran, Sarthak Kamat, Trevor Darrell, Koushil Sreenath, Jitendra Malik
Abstract
We cast real-world humanoid control as a next token prediction problem, akin to predicting the next word in language. Our model is a causal transformer trained via autoregressive prediction of sensorimotor trajectories. To account for the multi-modal nature of the data, we perform prediction in a modality-aligned way, and for each input token predict the next token from the same modality. This general formulation enables us to leverage data with missing modalities, like video trajectories without actions. We train our model on a collection of simulated trajectories coming from prior neural network policies, model-based controllers, motion capture data, and YouTube videos of humans. We show that our model enables a full-sized humanoid to walk in San Francisco zero-shot. Our model can transfer to the real world even when trained on only 27 hours of walking data, and can generalize to commands not seen during training like walking backward. These findings suggest a promising path toward learning challenging real-world control tasks by generative modeling of sensorimotor trajectories.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained TransformersLirui Wang, Xinlei Chen, Jialiang Zhao, Kaiming HeNeurIPS 2024 · 208 citations
- BAKU: An Efficient Transformer for Multi-Task Policy LearningSiddhant Haldar, Zhuoran Peng, Lerrel PintoNeurIPS 2024 · 120 citations
- 4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene CalibrationJiahui Zhang, Yurui Chen, Yueming Xu, Ze Huang et al.NeurIPS 2025 · 69 citations
- Omnigrasp: Grasping Diverse Objects with Simulated HumanoidsZhengyi Luo, Jinkun Cao, Sammy Christen, Alexander Winkler et al.NeurIPS 2024 · 66 citations
- From Experts to a Generalist: Toward General Whole-Body Control for Humanoid RobotsYuxuan Wang, Ming Yang, Gang Ding, Yu Zhang et al.NeurIPS 2025 · 36 citations
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
Related papers
- Humanoid Generative Pre-Training for Zero-Shot Motion TrackingZekun Qi, Xuchuan Chen, Jilong Wang, Chenghuai Lin et al.CVPR 2026 · 5 citations
- MetaMorph: Learning Universal Controllers with TransformersAgrim Gupta, Linxi Fan, Surya Ganguli, Li Fei-FeiICLR 2022 · 130 citations
- DrivingGPT: Unifying Driving World Modeling and Planning with Multi-Modal Autoregressive TransformersYuntao Chen, Yuqi Wang, Zhaoxiang ZhangICCV 2025 · 7 citations
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 950 citations
- End-to-End Language-Action Model for Humanoid Whole Body ControlYuxuan Wang, Haobin Jiang, Shiqing Yao, Ziluo Ding et al.CVPR 2026
