L-Verse: Bidirectional Generation Between Image and Text
Taehoon Kim, Gwangmo Song, Sihaeng Lee, Sangyun Kim, Yewon Seo, Soonyoung Lee, Seung Hwan Kim, Honglak Lee, Kyunghoon Bae
摘要
Far beyond learning long-range interactions of natural language, transformers are becoming the de-facto standard for many vision tasks with their power and scalability. Especially with cross-modal tasks between image and text, vector quantized variational autoencoders (VQ-VAEs) are widely used to make a raw RGB image into a sequence of feature vectors. To better leverage the correlation between image and text, we propose L-Verse, a novel architecture consisting of feature-augmented variational autoencoder (AugVAE) and bidirectional auto-regressive transformer (BiART) for image-to-text and text-to-image generation. Our AugVAE shows the state-of-the-art reconstruction performance on ImageNetlK validation set, along with the robustness to unseen images in the wild. Unlike other models, BiART can distinguish between image (or text) as a conditional reference and a generation target. L-Verse can be directly used for image-to-text or text-to-image generation without any finetuning or extra object detection framework. In quantitative and qualitative experiments, L-Verse shows impressive results against previous methods in both image-to-text and text-to-image generation on MS-COCO Captions. We furthermore assess the scalability of L-Verse architecture on Conceptual Captions and present the initial result of bidirectional vision-language representation learning on general domain.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Simplifying Multimodal Emotion Recognition with Single Eye Movement ModalityXu Yan, Li-Ming Zhao, Bao-Liang LuACM MM 2021 · 被引用 35 次
- Vector Quantized Wasserstein Auto-EncoderLong Tung Vuong, Trung Le, He Zhao, Chuanxia Zheng 等ICML 2023 · 被引用 24 次
- Locally Hierarchical Auto-Regressive Modeling for Image GenerationTackgeun You, Saehoon Kim, Chiheon Kim, Doyup Lee 等NeurIPS 2022 · 被引用 17 次
- CoBIT: A Contrastive Bi-directional Image-Text Generation ModelHaoxuan You, Mandy Guo, Zhecan Wang, Kai-Wei Chang 等ICLR 2024 · 被引用 15 次
- Do DALL-E and Flamingo Understand Each Other?Hang Li, Jindong Gu, Rajat Koner, Sahand Sharifzadeh 等ICCV 2023 · 被引用 14 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- CogView: Mastering Text-to-Image Generation via TransformersMing Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng 等NeurIPS 2021 · 被引用 1,026 次
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 被引用 992 次
相关 Paper
- MAGVLT: Masked Generative Vision-and-Language TransformerSungwoong Kim, Daejin Jo, Donghoon Lee, Jongmin KimCVPR 2023
- Language Quantized AutoEncoders: Towards Unsupervised Text-Image AlignmentHao Liu, Wilson Yan, Pieter AbbeelNeurIPS 2023 · 被引用 39 次
- Towards More Unified In-Context Visual UnderstandingDianmo Sheng, Dongdong Chen, Zhentao Tan, Qiankun Liu 等CVPR 2024 · 被引用 9 次
- TokenFlow: Unified Image Tokenizer for Multimodal Understanding and GenerationLiao Qu, Huichao Zhang, Yiheng Liu, Xu Wang 等CVPR 2025
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and ReconstructionSinan Du, Jiahao Guo, Bo Li, Shuhao Cui 等CVPR 2026 · 被引用 11 次
