Latent Normalizing Flows for Many-to-Many Cross-Domain Mappings
Shweta Mahajan, Iryna Gurevych, Stefan Roth
Abstract
Learned joint representations of images and text form the backbone of several important cross-domain tasks such as image captioning. Prior work mostly maps both domains into a common latent representation in a purely supervised fashion. This is rather restrictive, however, as the two domains follow distinct generative processes. Therefore, we propose a novel semi-supervised framework, which models shared information between domains and domain-specific information separately. The information shared between the domains is aligned with an invertible neural network. Our model integrates normalizing flow-based priors for the domain-specific information, which allows us to learn diverse many-to-many mappings between the two domains. We demonstrate the effectiveness of our model on diverse tasks, including image captioning and text-to-image synthesis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 48a22343-d552-403c-99d6-90efff2eed0fCited by top-tier papers11
- PortaSpeech: Portable and High-Quality Generative Text-to-SpeechYi Ren, Jinglin Liu, Zhou ZhaoNeurIPS 2021 · 97 citations
- Diverse Image Captioning with Context-Object Split Latent SpacesShweta Mahajan, Stefan RothNeurIPS 2020 · 47 citations
- Learning Distinct and Representative Modes for Image CaptioningQi Chen, Chaorui Deng, Qi WuNeurIPS 2022 · 27 citations
- Decoupling Global and Local Representations via Invertible Generative FlowsXuezhe Ma, Xiang Kong, Shanghang Zhang, Eduard H. HovyICLR 2021 · 24 citations
- Unity by Diversity: Improved Representation Learning for Multimodal VAEsThomas M. Sutter, Yang Meng, Andrea Agostini, Daphné Chopard et al.NeurIPS 2024 · 21 citations
Builds on1
Related papers
- AlignFlow: Cycle Consistent Learning from Multiple Domains via Normalizing FlowsAditya Grover, Christopher Chute, Rui Shu, Zhangjie Cao et al.AAAI 2020 · 72 citations
- Towards Unsupervised Image Captioning With Shared Multimodal EmbeddingsIro Laina, Christian Rupprecht, Nassir NavabICCV 2019 · 115 citations
- DINO: A Conditional Energy-Based GAN for Domain TranslationKonstantinos Vougioukas, Stavros Petridis, Maja PanticICLR 2021 · 8 citations
- Variational Distribution Learning for Unsupervised Text-to-Image GenerationMinsoo Kang, Doyup Lee, Jiseob Kim, Saehoon Kim et al.CVPR 2023
- Cross-Domain Correspondence Learning for Exemplar-Based Image TranslationPan Zhang, Bo Zhang, Dong Chen, Lu Yuan et al.CVPR 2020
