Optimus: Organizing Sentences via Pre-trained Modeling of a Latent Space
Chunyuan Li, Xiang Gao, Yuan Li, Baolin Peng, Xiujun Li, Yizhe Zhang, Jianfeng Gao
摘要
When trained effectively, the Variational Autoencoder (VAE) (Kingma and Welling, 2013; Bowman et al., 2016) can be both a powerful generative model and an effective representation learning framework for natural language. In this paper, we propose the first large-scale language VAE model OPTIMUS 1 . A universal latent embedding space for sentences is first pre-trained on large text corpus, and then fine-tuned for various language generation and understanding tasks. Compared with GPT-2, OPTIMUS enables guided language generation from an abstract level using the latent vectors. Compared with BERT, OPTIMUS can generalize better on low-resource language understanding tasks due to the smooth latent space structure. Extensive experimental results on a wide range of language tasks demonstrate the effectiveness of OPTIMUS. It achieves new state-of-the-art on VAE language modeling benchmarks. Encoder z x Decoder x h [CLS]
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper41
- Score-based Generative Modeling in Latent SpaceArash Vahdat, Karsten Kreis, Jan KautzNeurIPS 2021 · 被引用 903 次
- Any-to-Any Generation via Composable DiffusionZineng Tang, Ziyi Yang, Chenguang Zhu, Michael Zeng 等NeurIPS 2023 · 被引用 294 次
- Versatile Diffusion: Text, Images and Variations All in One Diffusion ModelXingqian Xu, Zhangyang Wang, Eric J. Zhang, Kai Wang 等ICCV 2023 · 被引用 265 次
- D2C: Diffusion-Decoding Models for Few-Shot Conditional GenerationAbhishek Sinha, Jiaming Song, Chenlin Meng, Stefano ErmonNeurIPS 2021 · 被引用 149 次
- Few-Shot Named Entity Recognition: An Empirical Baseline StudyJiaxin Huang, Chunyuan Li, Krishan Subudhi, Damien Jose 等EMNLP 2021 · 被引用 97 次
它引用的顶会 Paper4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- A Batch Normalized Inference Network Keeps the KL Vanishing AwayQile Zhu, Wei Bi, Xiaojiang Liu, Xiyao Ma 等ACL 2020 · 被引用 70 次
- Pre-train and Plug-in: Flexible Conditional Text Generation with Variational Auto-EncodersYu Duan, Canwen Xu, Jiaxin Pei, Jialong Han 等ACL 2020 · 被引用 31 次
相关 Paper
- Do sequence-to-sequence VAEs learn global features of sentences?Tom Bosc, Pascal VincentEMNLP 2020 · 被引用 5 次
- Improving Variational Autoencoders with Density Gap-based RegularizationJianfei Zhang, Jun Bai, Chenghua Lin, Yanmeng Wang 等NeurIPS 2022 · 被引用 11 次
- Scheduled Sampling in Vision-Language Pretraining with Decoupled Encoder-Decoder NetworkYehao Li, Yingwei Pan, Ting Yao, Jingwen Chen 等AAAI 2021 · 被引用 59 次
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- Topic-VQ-VAE: Leveraging Latent Codebooks for Flexible Topic-Guided Document GenerationYoungjoon Yoo, Jongwon ChoiAAAI 2024 · 被引用 8 次
