Zero-shot Synthesis with Group-Supervised Learning
Yunhao Ge, Sami Abu-El-Haija, Gan Xin, Laurent Itti
Abstract
Visual cognition of primates is superior to that of artificial neural networks in its ability to 'envision' a visual object, even a newly-introduced one, in different attributes including pose, position, color, texture, etc. To aid neural networks to envision objects with different attributes, we propose a family of objective functions, expressed on groups of examples, as a novel learning framework that we term Group-Supervised Learning (GSL). GSL decomposes inputs into a disentangled representation with swappable components that can be recombined to synthesize new samples, trained through similarity mining within groups of exemplars. For instance, images of red boats & blue cars can be decomposed and recombined to synthesize novel images of red cars. We describe a general class of datasets admissible by GSL. We propose an implementation based on auto-encoder, termed group-supervised zero-shot synthesis network (GZS-Net) trained with our learning framework, that can produce a high-quality red car even if no such example is witnessed during training. We test our model and learning framework on existing benchmarks, in addition to new dataset that we open-source. We qualitatively and quantitatively demonstrate that GZS-Net trained with GSL outperforms state-of-the-art methods
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eafa845c-c51f-46ea-baf3-d4a4fcf0065cCited by top-tier papers3
- Attribute-driven Disentangled Representation Learning for Multimodal RecommendationZhenyang Li, Fan Liu, Yinwei Wei, Zhiyong Cheng et al.ACM MM 2024 · 17 citations
- Graph-based Unsupervised Disentangled Representation Learning via Multimodal Large Language ModelsBaao Xie, Qiuyu Chen, Yunnan Wang, Zequn Zhang et al.NeurIPS 2024 · 15 citations
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja et al.ACM MM 2025 · 3 citations
Builds on1
Related papers
- Semantics Disentangling for Generalized Zero-Shot LearningZhi Chen, Yadan Luo, Ruihong Qiu, Sen Wang et al.ICCV 2021 · 143 citations
- Task-Independent Knowledge Makes for Transferable Representations for Generalized Zero-Shot LearningChaoqun Wang, Xuejin Chen, Shaobo Min, Xiaoyan Sun et al.AAAI 2021 · 22 citations
- MAGANet: Achieving Combinatorial Generalization by Modeling a Group ActionGeonho Hwang, Jaewoong Choi, Hyunsoo Cho, Myungjoo KangICML 2023 · 4 citations
- Generalized Zero-Shot Video Classification via Generative Adversarial NetworksMingyao Hong, Guorong Li, Xinfeng Zhang, Qingming HuangACM MM 2020 · 13 citations
- On Leveraging Variational Graph Embeddings for Open World Compositional Zero-Shot LearningMuhammad Umer Anwaar, Zhihui Pan, Martin KleinsteuberACM MM 2022 · 19 citations
