Learning Visual Generative Priors without Text
Shuailei Ma, Kecheng Zheng, Ying Wei, Wei Wu, Fan Lu, Yifei Zhang, Chen-Wei Xie, Biao Gong, Jiapeng Zhu, Yujun Shen
2025Year
3Top-tier citations
Abstract
4 Alibaba Group 5 HKUST https://ant-research.github.io/lumos (a) Text-to-Image Generation (b) Novel View Synthesis (c) Image-to-Video Generation Figure 1. Diverse downstream tasks of Lumos including (a) text-to-image generation, (b) novel view synthesis (left: input view, middle: random novel views, right: reconstruction Gaussian) and (c) image-to-video generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Reconstruction Alignment Improves Unified Multimodal ModelsJi Xie, Trevor Darrell, Luke Zettlemoyer, XuDong WangICLR 2026 · 52 citations
- Rethinking UMM Visual Generation: Masked Modeling for Efficient Image-Only Pre-trainingPeng Sun, Jun XIE, Tao LinCVPR 2026 · 1 citation
- Aligned Better, Listen Better for Audio-Visual Large Language ModelsYuxin Guo, Shuailei Ma, Shijie Ma, Xiaoyi Bao et al.ICLR 2025
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model PerspectiveHangjie Yuan, Weihua Chen, Jun Cen, Hu Yu et al.ICLR 2026 · 21 citations
- MOVIS: Enhancing Multi-Object Novel View Synthesis for Indoor ScenesRuijie Lu, Yixin Chen, Junfeng Ni, Baoxiong Jia et al.CVPR 2025
- UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible FeedbackPengwei Liu, Hangjie Yuan, Bo Dong, Jiazheng Xing et al.NeurIPS 2025 · 10 citations
- GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera ControlXuanchi Ren, Tianchang Shen, Jiahui Huang, Huan Ling et al.CVPR 2025
- LumiX: Structured and Coherent Text-to-Intrinsic GenerationXu Han, Biao Zhang, Xiangjun Tang, Xianzhi Li et al.CVPR 2026
