L-Verse: Bidirectional Generation Between Image and Text
Taehoon Kim, Gwangmo Song, Sihaeng Lee, Sangyun Kim, Yewon Seo, Soonyoung Lee, Seung Hwan Kim, Honglak Lee, Kyunghoon Bae
Abstract
Far beyond learning long-range interactions of natural language, transformers are becoming the de-facto standard for many vision tasks with their power and scalability. Especially with cross-modal tasks between image and text, vector quantized variational autoencoders (VQ-VAEs) are widely used to make a raw RGB image into a sequence of feature vectors. To better leverage the correlation between image and text, we propose L-Verse, a novel architecture consisting of feature-augmented variational autoencoder (AugVAE) and bidirectional auto-regressive transformer (BiART) for image-to-text and text-to-image generation. Our AugVAE shows the state-of-the-art reconstruction performance on ImageNetlK validation set, along with the robustness to unseen images in the wild. Unlike other models, BiART can distinguish between image (or text) as a conditional reference and a generation target. L-Verse can be directly used for image-to-text or text-to-image generation without any finetuning or extra object detection framework. In quantitative and qualitative experiments, L-Verse shows impressive results against previous methods in both image-to-text and text-to-image generation on MS-COCO Captions. We furthermore assess the scalability of L-Verse architecture on Conceptual Captions and present the initial result of bidirectional vision-language representation learning on general domain.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dcf1f66e-88ad-4ca5-ab34-8d7fcbdb2accCited by top-tier papers14
- Simplifying Multimodal Emotion Recognition with Single Eye Movement ModalityXu Yan, Li-Ming Zhao, Bao-Liang LuACM MM 2021 · 35 citations
- Vector Quantized Wasserstein Auto-EncoderLong Tung Vuong, Trung Le, He Zhao, Chuanxia Zheng et al.ICML 2023 · 24 citations
- Locally Hierarchical Auto-Regressive Modeling for Image GenerationTackgeun You, Saehoon Kim, Chiheon Kim, Doyup Lee et al.NeurIPS 2022 · 17 citations
- CoBIT: A Contrastive Bi-directional Image-Text Generation ModelHaoxuan You, Mandy Guo, Zhecan Wang, Kai-Wei Chang et al.ICLR 2024 · 15 citations
- Do DALL-E and Flamingo Understand Each Other?Hang Li, Jindong Gu, Rajat Koner, Sahand Sharifzadeh et al.ICCV 2023 · 14 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- CogView: Mastering Text-to-Image Generation via TransformersMing Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng et al.NeurIPS 2021 · 1,026 citations
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
Related papers
- MAGVLT: Masked Generative Vision-and-Language TransformerSungwoong Kim, Daejin Jo, Donghoon Lee, Jongmin KimCVPR 2023
- Language Quantized AutoEncoders: Towards Unsupervised Text-Image AlignmentHao Liu, Wilson Yan, Pieter AbbeelNeurIPS 2023 · 39 citations
- Towards More Unified In-Context Visual UnderstandingDianmo Sheng, Dongdong Chen, Zhentao Tan, Qiankun Liu et al.CVPR 2024 · 9 citations
- TokenFlow: Unified Image Tokenizer for Multimodal Understanding and GenerationLiao Qu, Huichao Zhang, Yiheng Liu, Xu Wang et al.CVPR 2025
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and ReconstructionSinan Du, Jiahao Guo, Bo Li, Shuhao Cui et al.CVPR 2026 · 11 citations
