Envisioning Beyond the Few: Disentangled Semantics and Primitives for Few-Shot Atypical Layout-to-Image Generation
Nan Bao, Yifan Zhao, Wenzhuang Wang, Jia Li
摘要
The layout-to-image (L2I) task enables fine-grained control over image generation via object categories and spatial layouts. However, existing L2I methods yield fragmented and distorted generations under few-shot atypical settings. We term this failure as representation fragmentation, arising from a granularity mismatch that entangles semantic identity with visual details. To address this issue, we propose a representation-driven framework that disentangles semantics from primitives for robust few-shot adaptation. Specifically, Semantic Anchoring aggregates categorical semantics into anchors for stable identity, while Primitive Imbuing models recomposable primitives for robust local detail modeling. Conceptual Steering further regulates optimization with a saliency-aware objective to preserve foreground semantic consistency. Extensive experiments demonstrate consistent improvements in the 5-shot regime over state-of-the-art L2I methods in both visual fidelity and alignment across diverse atypical domains. The source code is publicly available at https://github.com/iCVTEAM/DSP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- Few-shot Image Generation Using Discrete Content RepresentationYan Hong, Li Niu, Jianfu Zhang, Liqing ZhangACM MM 2022 · 被引用 11 次
- Semantic-Aware Generator and Low-level Feature Augmentation for Few-shot Image GenerationZhe Wang, Jiaoyan Guan, Mengping Yang, Ting Xiao 等ACM MM 2023 · 被引用 2 次
- Learned Spatial Representations for Few-shot Talking-Head SynthesisMoustafa Meshry, Saksham Suri, Larry S. Davis, Abhinav ShrivastavaICCV 2021 · 被引用 51 次
- SANSA: Unleashing the Hidden Semantics in SAM2 for Few-Shot SegmentationClaudia Cuttano, Gabriele Trivigno, Giuseppe Averta, Carlo MasoneNeurIPS 2025 · 被引用 9 次
- OntoAug: Rethinking Generative Data Augmentation via Ontology GuidanceShuo Wang, Zhichuan Wang, Jun LuoCVPR 2026
