DuoGen: Towards Autonomous Interleaved Multimodal Generation
Min Shi, Xiaohui Zeng, Jiannan Huang, Yin Cui, Francesco Ferroni, Jialuo Li, Zhaoshuo Li, Yogesh Balaji, Haoxiang Wang, Tsung-Yi Lin, Xiao Fu, Yue Zhao
2026Year
Abstract
A person sitting at a laboratory desk connected to a massive bioluminescent jellyfish with colorful cables, analog synthesizers, cinematic light Have the person in the first image stand by the pier in the last image, holding the smartphone in the second image with his hands. 1 2 3 Input Image Input Image Turn right. There is a large glass pyramid. Keep move forward. Move forward.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on24
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Generating Images with Multimodal Language ModelsJing Yu Koh, Daniel Fried, Russ SalakhutdinovNeurIPS 2023 · 403 citations
Related papers
- Guided Score identity Distillation for Data-Free One-Step Text-to-Image GenerationMingyuan Zhou, Zhendong Wang, Huangjie Zheng, Hai HuangICLR 2025
- CapHuman: Capture Your Moments in Parallel UniversesChao Liang, Fan Ma, Linchao Zhu, Yingying Deng et al.CVPR 2024
- Seeing in Extra Darkness Using a Deep-Red FlashJinhui Xiong, Jian Wang, Wolfgang Heidrich, Shree K. NayarCVPR 2021
- Mimir: Improving Video Diffusion Models for Precise Text UnderstandingShuai Tan, Biao Gong, Yutong Feng, Kecheng Zheng et al.CVPR 2025
- SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and TrainingJierun Chen, Dongting Hu, Xijie Huang, Huseyin Coskun et al.CVPR 2025
