Unpaired Image-to-Speech Synthesis With Multimodal Information Bottleneck
Shuang Ma, Daniel McDuff, Yale Song
Abstract
Deep generative models have led to significant advances in cross-modal generation such as text-to-image synthesis. Training these models typically requires paired data with direct correspondence between modalities. We introduce the novel problem of translating instances from one modality to another without paired data by leveraging an intermediate modality shared by the two other modalities. To demonstrate this, we take the problem of translating images to speech. In this case, one could leverage disjoint datasets with one shared modality, e.g., image-text pairs and text-speech pairs, with text as the shared modality. We call this problem ``skip-modal generation'' because the shared modality is skipped during the generation process. We propose a multimodal information bottleneck approach that learns the correspondence between modalities from unpaired data (image and speech) by leveraging the shared modality (text). We address fundamental challenges of skip-modal generation: 1) learning multimodal representations using a single model, 2) bridging the domain gap between two unrelated datasets, and 3) learning the correspondence between modalities from unpaired data. We show qualitative results on image-to-speech synthesis; this is the first time such results have been reported in the literature. We also show that our approach improves performance on traditional cross-modal generation, suggesting that it improves data efficiency in solving individual tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Self-Supervised MultiModal Versatile NetworksJean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelovic et al.NeurIPS 2020 · 423 citations
- Preserving Modality Structure Improves Multi-Modal LearningSirnam Swetha, Mamshad Nayeem Rizve, Nina Shvetsova, Hilde Kuehne et al.ICCV 2023 · 15 citations
- Everything at Once - Multi-modal Fusion Transformer for Video RetrievalNina Shvetsova, Brian Chen, Andrew Rouditchenko, Samuel Thomas et al.CVPR 2022 · 4 citations
- Learning Shared Representations from Unpaired DataAmitai Yacobi, Nir Ben-Ari, Ronen Talmon, Uri ShahamNeurIPS 2025 · 3 citations
- Farewell to Mutual Information: Variational Distillation for Cross-Modal Person Re-IdentificationXudong Tian, Zhizhong Zhang, Shaohui Lin, Yanyun Qu et al.CVPR 2021
Related papers
- Bridging the Gap between Synthetic and Authentic Images for Multimodal Machine TranslationWenyu Guo, Qingkai Fang, Dong Yu, Yang FengEMNLP 2023 · 5 citations
- UNISON: Unpaired Cross-Lingual Image CaptioningJiahui Gao, Yi Zhou, Philip L. H. Yu, Shafiq R. Joty et al.AAAI 2022 · 18 citations
- Disentangled Cross-Modal Representation Learning with Enhanced Mutual SupervisionLu Gao, Wenlan Chen, Daoyuan Wang, Fei Guo et al.NeurIPS 2025 · 5 citations
- T-Modules: Translation Modules for Zero-Shot Cross-Modal Machine TranslationPaul-Ambroise Duquenne, Hongyu Gong, Benoît Sagot, Holger SchwenkEMNLP 2022 · 8 citations
- Cross-Modal Knowledge Distillation without Paired Data: Theoretical Foundation and AlgorithmT. K Tran, Duc Chu Anh, Quang Hung Pham, Phi Le Nguyen et al.ICML 2026
