Disentanglement via Latent Quantization
Kyle Hsu, William Dorrell, James C. R. Whittington, Jiajun Wu, Chelsea Finn
摘要
In disentangled representation learning, a model is asked to tease apart a dataset's underlying sources of variation and represent them independently of one another. Since the model is provided with no ground truth information about these sources, inductive biases take a paramount role in enabling disentanglement. In this work, we construct an inductive bias towards encoding to and decoding from an organized latent space. Concretely, we do this by (i) quantizing the latent space into discrete code vectors with a separate learnable scalar codebook per dimension and (ii) applying strong model regularization via an unusually high weight decay. Intuitively, the latent space design forces the encoder to combinatorially construct codes from a small number of distinct scalar values, which in turn enables the decoder to assign a consistent meaning to each value. Regularization then serves to drive the model towards this parsimonious strategy. We demonstrate the broad applicability of this approach by adding it to both basic data-reconstructing (vanilla autoencoder) and latent-reconstructing (InfoGAN) generative models. For reliable evaluation, we also propose InfoMEC, a new set of metrics for disentanglement that is cohesively grounded in information theory and fixes well-established shortcomings in previous metrics. Together with regularization, latent quantization dramatically improves the modularity and explicitness of learned representations on a representative suite of benchmark datasets. In particular, our quantized-latent autoencoder (QLAE) consistently outperforms strong methods from prior work in these key disentanglement properties without compromising data reconstruction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Finite Scalar Quantization: VQ-VAE Made SimpleFabian Mentzer, David Minnen, Eirikur Agustsson, Michael TschannenICLR 2024 · 被引用 442 次
- Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)Usha Bhalla, Alex Oesterling, Suraj Srinivas, Flávio P. Calmon 等NeurIPS 2024 · 被引用 146 次
- Selective Visual Representations Improve Convergence and Generalization for Embodied AIAinaz Eftekhar, Kuo-Hao Zeng, Jiafei Duan, Ali Farhadi 等ICLR 2024 · 被引用 28 次
- Tripod: Three Complementary Inductive Biases for Disentangled Representation LearningKyle Hsu, Jubayer Ibn Hamid, Kaylee Burns, Chelsea Finn 等ICML 2024 · 被引用 13 次
- Language-Informed Visual Concept LearningSharon Lee, Yunzhi Zhang, Shangzhe Wu, Jiajun WuICLR 2024 · 被引用 13 次
它引用的顶会 Paper21
- vq-wav2vec: Self-Supervised Learning of Discrete Speech RepresentationsAlexei Baevski, Steffen Schneider, Michael AuliICLR 2020 · 被引用 730 次
- Weakly-Supervised Disentanglement Without CompromisesFrancesco Locatello, Ben Poole, Gunnar Rätsch, Bernhard Schölkopf 等ICML 2020 · 被引用 361 次
- A Theory of Usable Information under Computational ConstraintsYilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart 等ICLR 2020 · 被引用 211 次
- Towards Nonlinear Disentanglement in Natural Data with Temporal Sparse CodingDavid A. Klindt, Lukas Schott, Yash Sharma, Ivan Ustyuzhaninov 等ICLR 2021 · 被引用 156 次
- Weakly Supervised Disentanglement with GuaranteesRui Shu, Yining Chen, Abhishek Kumar, Stefano Ermon 等ICLR 2020 · 被引用 148 次
相关 Paper
- Theory and Evaluation Metrics for Learning Disentangled RepresentationsKien Do, Truyen TranICLR 2020 · 被引用 107 次
- IB-GAN: Disentangled Representation Learning with Information Bottleneck Generative Adversarial NetworksInsu Jeon, Wonkwang Lee, Myeongjang Pyeon, Gunhee KimAAAI 2021 · 被引用 47 次
- Information-theoretic Generalization Analysis for VQ-VAEs: A Role of Latent VariablesFutoshi Futami, Masahiro FujisawaNeurIPS 2025 · 被引用 1 次
- Orthogonality-Enforced Latent Space in Autoencoders: An Approach to Learning Disentangled RepresentationsJaehoon Cha, Jeyan ThiyagalingamICML 2023 · 被引用 1 次
- A Bayesian Nonparametric Framework For Learning Disentangled RepresentationsVaishnavi Patil, Siddhi Patil, Matthew Evanusa, Amit Kumar Kundu 等ICLR 2026
