Hölder++: Improving Quality-Coherence Trade-off in Multimodal VAEs
Huyen Vo, María Martínez-García, Isabel Valera
摘要
Existing approaches for multimodal variational autoencoders (VAEs) face a trade-off between generative quality and coherence—i.e., they struggle to generate realistic and diverse samples that, at the same time, are semantically consistent across modalities. A recent work shows that using a simple approximation to Hölder pooling as an aggregation method improves coherence over the SOTA MMVAE+, despite assuming a single shared representation across all modalities. Yet, it slightly compromises sample diversity. Inspired by this insight, we propose Hölder++, a novel multimodal VAE that improves the generative quality-coherence trade-off through: (i) the first implementation of Hölder pooling without any approximation for multimodal VAEs; (ii) an extended architecture that models distinct shared and private (i.e., modality-specific) representations (Hölder+); and (iii) hierarchical inference that further enhances the disentanglement between the shared and private representations (Hölder++). Our experiments corroborate that Hölder++ consistently improves the generative quality-coherence trade-off, yields more structured latent spaces, and learns shared representations that are informative for downstream tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 被引用 1,141 次
- Self-Supervised Learning with Data Augmentations Provably Isolates Content from StyleJulius von Kügelgen, Yash Sharma, Luigi Gresele, Wieland Brendel 等NeurIPS 2021 · 被引用 421 次
- Generalized Multimodal ELBOThomas M. Sutter, Imant Daunhawer, Julia E. VogtICLR 2021 · 被引用 130 次
- Hierarchical VAEs Know What They Don't KnowJakob Drachmann Havtorn, Jes Frellsen, Søren Hauberg, Lars MaaløeICML 2021 · 被引用 87 次
- On the Limitations of Multimodal VAEsImant Daunhawer, Thomas M. Sutter, Kieran Chin-Cheong, Emanuele Palumbo 等ICLR 2022 · 被引用 50 次
相关 Paper
- MMVAE+: Enhancing the Generative Quality of Multimodal VAEs without CompromisesEmanuele Palumbo, Imant Daunhawer, Julia E. VogtICLR 2023
- Unity by Diversity: Improved Representation Learning for Multimodal VAEsThomas M. Sutter, Yang Meng, Andrea Agostini, Daphné Chopard 等NeurIPS 2024 · 被引用 21 次
- Disentanglement of Variations with Multimodal Generative ModelingYijie Zhang, Yiyang Shen, Weiran WangICLR 2026 · 被引用 6 次
- ShaLa: Multimodal Shared Latent Generative ModellingJiali Cui, Yan-Ying Chen, Yanxia Zhang, Matthew KlenkAAAI 2026
- Efficient Modality Translation via Arbitrary Conditioning and Wasserstein RegularizationTomás Tokár, Scott SannerAAAI 2026
