Generating Multi-Image Synthetic Data for Text-to-Image Customization
Nupur Kumari, Xi Yin, Jun-Yan Zhu, Ishan Misra, Samaneh Azadi
Abstract
Customization of text-to-image models enables users to insert new concepts or objects and generate them in unseen settings. Existing methods either rely on comparatively expensive test-time optimization or train encoders on singleimage datasets without multi-image supervision, which can limit image quality. We propose a simple approach to address these challenges. We first leverage existing text-to-image models and 3D datasets to create a high-quality Synthetic Customization Dataset (SynCD) consisting of multiple images of the same object in different lighting, backgrounds, and poses. Using this dataset, we train an encoder-based model that incorporates fine-grained visual details from reference images via a shared attention mechanism. Finally, we propose an inference technique that normalizes text and image guidance vectors to mitigate overexposure issues in sampled images. Through extensive experiments, we show that our encoder-based model, trained on SynCD, and with the proposed inference algorithm, improves upon existing encoder-based methods on standard customization benchmarks. Please find the code and data at our website.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Scaling Group Inference for Diverse and High-Quality GenerationGaurav Parmar, Or Patashnik, Daniil Ostashev, Kuan-Chieh Wang et al.ICLR 2026 · 14 citations
- Learning an Image Editing Model without Image Editing PairsNupur Kumari, Sheng-Yu Wang, Nanxuan Zhao, Yotam Nitzan et al.ICLR 2026 · 14 citations
- UP2You: Fast Reconstruction of Yourself from Unconstrained Photo CollectionsZeyu Cai, Ziyang Li, Xiaoben Li, Boqian Li et al.ICLR 2026 · 11 citations
- IC-Custom: Diverse Image Customization via In-Context LearningYaowei Li, Xiaoyu Li, Zhaoyang Zhang, Yuxuan Bian et al.ICLR 2026 · 10 citations
- Dimension-Reduction Attack! Video Generative Models are Experts on Controllable Image SynthesisHengyuan Cao, Yutong Feng, Biao Gong, Yijing Tian et al.NeurIPS 2025 · 6 citations
Builds on50
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
Related papers
- CustomNet: Object Customization with Variable-Viewpoints in Text-to-Image Diffusion ModelsZiyang Yuan, Mingdeng Cao, Xintao Wang, Zhongang Qi et al.ACM MM 2024 · 10 citations
- MC^2: Multi-concept Guidance for Customized Multi-concept GenerationJiaxiu Jiang, Yabo Zhang, Kailai Feng, Xiaohe Wu et al.CVPR 2025
- Encoder-based Domain Tuning for Fast Personalization of Text-to-Image ModelsRinon Gal, Moab Arar, Yuval Atzmon, Amit H. Bermano et al.SIGGRAPH 2023 · 154 citations
- SPAD: Spatially Aware Multi-View DiffusersYash Kant, Aliaksandr Siarohin, Ziyi Wu, Michael Vasilkovsky et al.CVPR 2024
- Visual Persona: Foundation Model for Full-Body Human CustomizationJisu Nam, Soowon Son, Zhan Xu, Jing Shi et al.CVPR 2025
