PatchVAE: Learning Local Latent Codes for Recognition
Kamal Gupta, Saurabh Singh, Abhinav Shrivastava
Abstract
Unsupervised representation learning holds the promise of exploiting large amounts of unlabeled data to learn general representations. A promising technique for unsupervised learning is the framework of Variational Auto-encoders (VAEs). However, unsupervised representations learned by VAEs are significantly outperformed by those learned by supervised learning for recognition. Our hypothesis is that to learn useful representations for recognition the model needs to be encouraged to learn about repeating and consistent patterns in data. Drawing inspiration from the mid-level representation discovery work, we propose PatchVAE, that reasons about images at patch level. Our key contribution is a bottleneck formulation that encourages mid-level style representations in the VAE framework. Our experiments demonstrate that representations learned by our method perform much better on the recognition tasks compared to those learned by vanilla VAEs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- LayoutTransformer: Layout Generation and Completion with Self-attentionKamal Gupta, Justin Lazarow, Alessandro Achille, Larry Davis et al.ICCV 2021 · 184 citations
- Hierarchical Patch VAE-GAN: Generating Diverse Videos from a Single SampleShir Gur, Sagie Benaim, Lior WolfNeurIPS 2020 · 84 citations
- GANSeg: Learning to Segment by Unsupervised Hierarchical Image GenerationXingzhe He, Bastian Wandt, Helge RhodinCVPR 2022 · 19 citations
- Unsupervised Causal Binary Concepts Discovery with VAE for Black-Box Model ExplanationThien Q. Tran, Kazuto Fukuchi, Youhei Akimoto, Jun SakumaAAAI 2022 · 11 citations
- PatchGame: Learning to Signal Mid-level Patches in Referential GamesKamal Gupta, Gowthami Somepalli, Anubhav Gupta, Vinoj Yasanga Jayasundara Magalle Hewa et al.NeurIPS 2021 · 4 citations
Related papers
- Unsupervised Part Discovery via Descriptor-Based Masked Image Restoration with Optimized ConstraintsJiahao Xia, Yike Wu, Wenjian Huang, Jianguo Zhang et al.ICCV 2025 · 1 citation
- Self-Guided Masked AutoencoderJeongwoo Shin, Inseo Lee, Junho Lee, Joonseok LeeNeurIPS 2024 · 18 citations
- Adversarial Disentanglement with Grouped ObservationsJózsef NémethAAAI 2020 · 8 citations
- Structure by Architecture: Structured Representations without RegularizationFelix Leeb, Giulia Lanzillotta, Yashas Annadani, Michel Besserve et al.ICLR 2023 · 1 citation
- Understanding Masked Autoencoders via Hierarchical Latent Variable ModelsLingjing Kong, Martin Q. Ma, Guangyi Chen, Eric P. Xing et al.CVPR 2023
