Conditional Generative Modeling via Learning the Latent Space
Sameera Ramasinghe, Kanchana Nisal Ranasinghe, Salman H. Khan, Nick Barnes, Stephen Gould
摘要
Although deep learning has achieved appealing results on several machine learning tasks, most of the models are deterministic at inference, limiting their application to single-modal settings. We propose a novel general-purpose framework for conditional generation in multimodal spaces, that uses latent variables to model generalizable learning patterns while minimizing a family of regression cost functions. At inference, the latent variables are optimized to find optimal solutions corresponding to multiple output modes. Compared to existing generative solutions, in multimodal spaces, our approach demonstrates faster and stable convergence, and can learn better representations for downstream tasks. Importantly, it provides a simple generic model that can beat highly engineered pipelines tailored using domain expertise on a variety of tasks, while generating diverse outputs. Our codes will be released.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Accelerating Text-to-Image Editing via Cache-Enabled Sparse Diffusion InferenceZihao Yu, Haoyang Li, Fangcheng Fu, Xupeng Miao 等AAAI 2024 · 被引用 19 次
- Rethinking conditional GAN training: An approach using geometrically structured latent manifoldsSameera Ramasinghe, Moshiur R. Farazi, Salman H. Khan, Nick Barnes 等NeurIPS 2021 · 被引用 11 次
它引用的顶会 Paper3
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud 等ICLR 2020 · 被引用 643 次
- Smoothness and Stability in GANsCasey Chu, Kentaro Minami, Kenji FukumizuICLR 2020 · 被引用 65 次
- Understanding the Limitations of Conditional Generative ModelsEthan Fetaya, Jörn-Henrik Jacobsen, Will Grathwohl, Richard S. ZemelICLR 2020 · 被引用 65 次
相关 Paper
- Executing your Commands via Motion Diffusion in Latent SpaceXin Chen, Biao Jiang, Wen Liu, Zilong Huang 等CVPR 2023
- Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent RepresentationZibo Zhao, Wen Liu, Xin Chen, Xianfang Zeng 等NeurIPS 2023 · 被引用 279 次
- Generalized Multimodal ELBOThomas M. Sutter, Imant Daunhawer, Julia E. VogtICLR 2021 · 被引用 130 次
- ShaLa: Multimodal Shared Latent Generative ModellingJiali Cui, Yan-Ying Chen, Yanxia Zhang, Matthew KlenkAAAI 2026
- Multilinear Latent Conditioning for Generating Unseen Attribute CombinationsMarkos Georgopoulos, Grigorios Chrysos, Maja Pantic, Yannis PanagakisICML 2020 · 被引用 18 次
