Unpaired Image-to-Speech Synthesis With Multimodal Information Bottleneck
Shuang Ma, Daniel McDuff, Yale Song
摘要
Deep generative models have led to significant advances in cross-modal generation such as text-to-image synthesis. Training these models typically requires paired data with direct correspondence between modalities. We introduce the novel problem of translating instances from one modality to another without paired data by leveraging an intermediate modality shared by the two other modalities. To demonstrate this, we take the problem of translating images to speech. In this case, one could leverage disjoint datasets with one shared modality, e.g., image-text pairs and text-speech pairs, with text as the shared modality. We call this problem ``skip-modal generation'' because the shared modality is skipped during the generation process. We propose a multimodal information bottleneck approach that learns the correspondence between modalities from unpaired data (image and speech) by leveraging the shared modality (text). We address fundamental challenges of skip-modal generation: 1) learning multimodal representations using a single model, 2) bridging the domain gap between two unrelated datasets, and 3) learning the correspondence between modalities from unpaired data. We show qualitative results on image-to-speech synthesis; this is the first time such results have been reported in the literature. We also show that our approach improves performance on traditional cross-modal generation, suggesting that it improves data efficiency in solving individual tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Self-Supervised MultiModal Versatile NetworksJean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelovic 等NeurIPS 2020 · 被引用 423 次
- Preserving Modality Structure Improves Multi-Modal LearningSirnam Swetha, Mamshad Nayeem Rizve, Nina Shvetsova, Hilde Kuehne 等ICCV 2023 · 被引用 15 次
- Everything at Once - Multi-modal Fusion Transformer for Video RetrievalNina Shvetsova, Brian Chen, Andrew Rouditchenko, Samuel Thomas 等CVPR 2022 · 被引用 4 次
- Learning Shared Representations from Unpaired DataAmitai Yacobi, Nir Ben-Ari, Ronen Talmon, Uri ShahamNeurIPS 2025 · 被引用 3 次
- Farewell to Mutual Information: Variational Distillation for Cross-Modal Person Re-IdentificationXudong Tian, Zhizhong Zhang, Shaohui Lin, Yanyun Qu 等CVPR 2021
相关 Paper
- Bridging the Gap between Synthetic and Authentic Images for Multimodal Machine TranslationWenyu Guo, Qingkai Fang, Dong Yu, Yang FengEMNLP 2023 · 被引用 5 次
- UNISON: Unpaired Cross-Lingual Image CaptioningJiahui Gao, Yi Zhou, Philip L. H. Yu, Shafiq R. Joty 等AAAI 2022 · 被引用 18 次
- Disentangled Cross-Modal Representation Learning with Enhanced Mutual SupervisionLu Gao, Wenlan Chen, Daoyuan Wang, Fei Guo 等NeurIPS 2025 · 被引用 5 次
- T-Modules: Translation Modules for Zero-Shot Cross-Modal Machine TranslationPaul-Ambroise Duquenne, Hongyu Gong, Benoît Sagot, Holger SchwenkEMNLP 2022 · 被引用 8 次
- Cross-Modal Knowledge Distillation without Paired Data: Theoretical Foundation and AlgorithmT. K Tran, Duc Chu Anh, Quang Hung Pham, Phi Le Nguyen 等ICML 2026
