AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction
Junhao Cheng, Yuying Ge, Yixiao Ge, Jing Liao, Ying Shan
Abstract
Recent advancements in image and video synthesis have opened up new promise in generative games. One particularly intriguing application is transforming characters from anime films into interactive, playable entities. This allows players to immerse themselves in the dynamic anime world as their favorite characters for life simulation through language instructions. Such games are defined as infinite game since they eliminate predetermined boundaries and fixed gameplay rules, where players can interact with the game world through open-ended language and experience ever-evolving storylines and environments. Recently, a pioneering approach for infinite anime life simulation employs large language models (LLMs) to translate multi-turn text dialogues into language instructions for image generation. However, it neglects historical visual context, leading to inconsistent gameplay. Furthermore, it only generates static images, failing to incorporate the dynamics necessary for an engaging gaming experience. In this work, we propose AnimeGamer, which is built upon Multimodal Large Language Models (MLLMs) to generate each game state, including dynamic animation shots that depict character movements and updates to character states, as illustrated in Figure 1. We introduce novel action-aware multimodal representations to represent animation shots, which can be decoded into high-quality video clips using a video diffusion model. By taking historical animation shot representations as context and predicting subsequent representations, AnimeGamer can generate games with contextual consistency and satisfactory dynamics. Extensive evaluations using both automated metrics and human evaluations demonstrate that AnimeGamer outperforms existing methods in various aspects of the gaming experience. Codes and checkpoints are available at https://github.com/TencentARC/AnimeGamer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c377e16d-5149-4466-8053-96cedf5c9532Cited by top-tier papers3
- PlayerOne: Egocentric World SimulatorYuanpeng Tu, Hao Luo, Xi Chen, Xiang Bai et al.NeurIPS 2025 · 20 citations
- Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPOJunhao Cheng, Liang Hou, Xin Tao, Jing LiaoCVPR 2026 · 6 citations
- Beyond the Golden Data: Resolving the Motion-Vision Quality Dilemma via Timestep Selective TrainingXiangyang Luo, Qingyu Li, Yuming Li, Guanbo Huang et al.CVPR 2026 · 3 citations
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
Related papers
- Unbounded: A Generative Infinite Game of Character Life SimulationJialu Li, Yuanzhen Li, Neal Wadhwa, Yael Pritch et al.ICLR 2025
- Role-Aware Virtual Agents for Navigational Interaction guided by a Multimodal Large Language ModelMinyoung Kim, Changyang Li, Cuong Nguyen, Lap-Fai YuSIGGRAPH 2026
- Text-Guided Synthesis of Crowd AnimationXuebo Ji, Zherong Pan, Xifeng Gao, Jia PanSIGGRAPH 2024 · 4 citations
- Instilling an Active Mind in Avatars via Cognitive SimulationJianwen Jiang, Weihong Zeng, Zerong Zheng, Jiaqi Yang et al.ICLR 2026 · 26 citations
- LGTM: Local-to-Global Text-Driven Human Motion Diffusion ModelHaowen Sun, Ruikun Zheng, Haibin Huang, Chongyang Ma et al.SIGGRAPH 2024 · 12 citations
