Say Cheese! Detail-Preserving Portrait Collection Generation via Natural Language Edits
Zelong Sun, Jiahui Wu, Ying Ba, Dong Jing, Zhiwu Lu
Abstract
As social media platforms proliferate, users increasingly demand intuitive ways to create diverse, high-quality portrait collections. In this work, we introduce Portrait Collection Generation (PCG), a novel task that generates coherent portrait collections by editing a reference portrait image through natural language instructions. This task poses two unique challenges to existing methods: (1) complex multi-attribute modifications such as pose, spatial layout, and camera viewpoint; and (2) high-fidelity detail preservation including identity, clothing, and accessories. To address these challenges, we propose CHEESE , the first large-scale PCG dataset containing 24K portrait collections and 573K samples with high-quality modification text annotations, constructed through an Large Vison-Language Model-based pipeline with inversion-based verification. We further propose SCheese , a framework that combines text-guided generation with hierarchical identity and detail preservation. SCheese employs adaptive feature fusion mechanism to maintain identity consistency, and ConsistencyNet to inject fine-grained features for detail consistency. Comprehensive experiments validate the effectiveness of CHEESE in advancing PCG, with SCheese achieving state-of-the-art performance in handling complex edits with identity and fine-grained details consistency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 83f3f06b-9c96-4c9e-a246-36f53ab69df6Cited by top-tier papers2
- PortraitRL: Reinforcement Learning for Personalized Portrait Pose Transfer with Multi-Objective Reward ModelingJiahui Wu, Zelong Sun, Yanbiao Ma, Zhiwu LuICML 2026
- Pareto-Guided Optimal Transport for Multi-Reward AlignmentYing Ba, Tianyu Zhang, Mohan Zhou, Yalong Bai et al.ICML 2026
Builds on24
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- EchoShot: Multi-Shot Portrait Video GenerationJiahao Wang, Hualian Sheng, Sijia Cai, Weizhan Zhang et al.NeurIPS 2025 · 30 citations
- HairCLIP: Design Your Hair by Text and Reference ImageTianyi Wei, Dongdong Chen, Wenbo Zhou, Jing Liao et al.CVPR 2022 · 94 citations
- Visual Persona: Foundation Model for Full-Body Human CustomizationJisu Nam, Soowon Son, Zhan Xu, Jing Shi et al.CVPR 2025
- ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video GenerationMingyang Wu, Ashirbad Mishra, Soumik Dey, Shuo Xing et al.CVPR 2026 · 8 citations
- Diverse Person: Customize Your Own Dataset for Text-Based Person SearchZifan Song, Guosheng Hu, Cairong ZhaoAAAI 2024
