DuoGen: Towards Autonomous Interleaved Multimodal Generation
Min Shi, Xiaohui Zeng, Jiannan Huang, Yin Cui, Francesco Ferroni, Jialuo Li, Zhaoshuo Li, Yogesh Balaji, Haoxiang Wang, Tsung-Yi Lin, Xiao Fu, Yue Zhao
2026年份
摘要
A person sitting at a laboratory desk connected to a massive bioluminescent jellyfish with colorful cables, analog synthesizers, cinematic light Have the person in the first image stand by the pier in the last image, holding the smartphone in the second image with his hands. 1 2 3 Input Image Input Image Turn right. There is a large glass pyramid. Keep move forward. Move forward.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Generating Images with Multimodal Language ModelsJing Yu Koh, Daniel Fried, Russ SalakhutdinovNeurIPS 2023 · 被引用 403 次
相关 Paper
- Guided Score identity Distillation for Data-Free One-Step Text-to-Image GenerationMingyuan Zhou, Zhendong Wang, Huangjie Zheng, Hai HuangICLR 2025
- CapHuman: Capture Your Moments in Parallel UniversesChao Liang, Fan Ma, Linchao Zhu, Yingying Deng 等CVPR 2024
- Seeing in Extra Darkness Using a Deep-Red FlashJinhui Xiong, Jian Wang, Wolfgang Heidrich, Shree K. NayarCVPR 2021
- Mimir: Improving Video Diffusion Models for Precise Text UnderstandingShuai Tan, Biao Gong, Yutong Feng, Kecheng Zheng 等CVPR 2025
- SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and TrainingJierun Chen, Dongting Hu, Xijie Huang, Huseyin Coskun 等CVPR 2025
