Latent Normalizing Flows for Many-to-Many Cross-Domain Mappings
Shweta Mahajan, Iryna Gurevych, Stefan Roth
摘要
Learned joint representations of images and text form the backbone of several important cross-domain tasks such as image captioning. Prior work mostly maps both domains into a common latent representation in a purely supervised fashion. This is rather restrictive, however, as the two domains follow distinct generative processes. Therefore, we propose a novel semi-supervised framework, which models shared information between domains and domain-specific information separately. The information shared between the domains is aligned with an invertible neural network. Our model integrates normalizing flow-based priors for the domain-specific information, which allows us to learn diverse many-to-many mappings between the two domains. We demonstrate the effectiveness of our model on diverse tasks, including image captioning and text-to-image synthesis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- PortaSpeech: Portable and High-Quality Generative Text-to-SpeechYi Ren, Jinglin Liu, Zhou ZhaoNeurIPS 2021 · 被引用 97 次
- Diverse Image Captioning with Context-Object Split Latent SpacesShweta Mahajan, Stefan RothNeurIPS 2020 · 被引用 47 次
- Learning Distinct and Representative Modes for Image CaptioningQi Chen, Chaorui Deng, Qi WuNeurIPS 2022 · 被引用 27 次
- Decoupling Global and Local Representations via Invertible Generative FlowsXuezhe Ma, Xiang Kong, Shanghang Zhang, Eduard H. HovyICLR 2021 · 被引用 24 次
- Unity by Diversity: Improved Representation Learning for Multimodal VAEsThomas M. Sutter, Yang Meng, Andrea Agostini, Daphné Chopard 等NeurIPS 2024 · 被引用 21 次
它引用的顶会 Paper1
相关 Paper
- AlignFlow: Cycle Consistent Learning from Multiple Domains via Normalizing FlowsAditya Grover, Christopher Chute, Rui Shu, Zhangjie Cao 等AAAI 2020 · 被引用 72 次
- Towards Unsupervised Image Captioning With Shared Multimodal EmbeddingsIro Laina, Christian Rupprecht, Nassir NavabICCV 2019 · 被引用 115 次
- DINO: A Conditional Energy-Based GAN for Domain TranslationKonstantinos Vougioukas, Stavros Petridis, Maja PanticICLR 2021 · 被引用 8 次
- Variational Distribution Learning for Unsupervised Text-to-Image GenerationMinsoo Kang, Doyup Lee, Jiseob Kim, Saehoon Kim 等CVPR 2023
- Cross-Domain Correspondence Learning for Exemplar-Based Image TranslationPan Zhang, Bo Zhang, Dong Chen, Lu Yuan 等CVPR 2020
