Show-o2: Improved Native Unified Multimodal Models
Jinheng Xie, Zhenheng Yang, Mike Zheng Shou
摘要
This paper presents improved native unified multimodal models, i.e., Show-o2, that leverage autoregressive modeling and flow matching. Built upon a 3D causal variational autoencoder space, unified visual representations are constructed through a dual-path of spatial (-temporal) fusion, enabling scalability across image and video modalities while ensuring effective multimodal understanding and generation. Based on a language model, autoregressive modeling and flow matching are natively applied to the language head and flow head, respectively, to facilitate text token prediction and image/video generation. A two-stage training recipe is designed to effectively learn and scale to larger models. The resulting Show-o2 models demonstrate versatility in handling a wide range of multimodal understanding and generation tasks across diverse modalities, including text, images, and videos. Code and models are released at https://github.com/showlab/Show-o.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper78
- Motus: A Unified Latent Action World ModelHongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang 等CVPR 2026 · 被引用 271 次
- UniVideo: Unified Understanding, Generation, and Editing for VideosCong Wei, Quande Liu, Zixuan Ye, Qiulin Wang 等ICLR 2026 · 被引用 90 次
- Latent Visual ReasoningBangzheng Li, Ximeng Sun, Jiang Liu, Ze Wang 等ICLR 2026 · 被引用 80 次
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at ScaleChunrui Han, Guopeng Li, Jingwei Wu, Quan Sun 等ICLR 2026 · 被引用 58 次
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and GenerationHan Li, Xinyu Peng, Yaoming Wang, Zelin Peng 等CVPR 2026 · 被引用 47 次
它引用的顶会 Paper60
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
相关 Paper
- Show-o: One Single Transformer to Unify Multimodal Understanding and GenerationJinheng Xie, Weijia Mao, Zechen Bai, David Junhao Zhang 等ICLR 2025
- FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal VelocitiesJin Wang, Yao Lai, Aoxue Li, Shifeng Zhang 等NeurIPS 2025 · 被引用 45 次
- VILA-U: a Unified Foundation Model Integrating Visual Understanding and GenerationYecheng Wu, Zhuoyang Zhang, Junyu Chen, Haotian Tang 等ICLR 2025
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal ModelsZhiheng Liu, Weiming Ren, Haozhe Liu, Zijian Zhou 等CVPR 2026 · 被引用 36 次
- FlowAR: Scale-wise Autoregressive Image Generation Meets Flow MatchingSucheng Ren, Qihang Yu, Ju He, Xiaohui Shen 等ICML 2025
