Generating Multi-Image Synthetic Data for Text-to-Image Customization
Nupur Kumari, Xi Yin, Jun-Yan Zhu, Ishan Misra, Samaneh Azadi
摘要
Customization of text-to-image models enables users to insert new concepts or objects and generate them in unseen settings. Existing methods either rely on comparatively expensive test-time optimization or train encoders on singleimage datasets without multi-image supervision, which can limit image quality. We propose a simple approach to address these challenges. We first leverage existing text-to-image models and 3D datasets to create a high-quality Synthetic Customization Dataset (SynCD) consisting of multiple images of the same object in different lighting, backgrounds, and poses. Using this dataset, we train an encoder-based model that incorporates fine-grained visual details from reference images via a shared attention mechanism. Finally, we propose an inference technique that normalizes text and image guidance vectors to mitigate overexposure issues in sampled images. Through extensive experiments, we show that our encoder-based model, trained on SynCD, and with the proposed inference algorithm, improves upon existing encoder-based methods on standard customization benchmarks. Please find the code and data at our website.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Scaling Group Inference for Diverse and High-Quality GenerationGaurav Parmar, Or Patashnik, Daniil Ostashev, Kuan-Chieh Wang 等ICLR 2026 · 被引用 14 次
- Learning an Image Editing Model without Image Editing PairsNupur Kumari, Sheng-Yu Wang, Nanxuan Zhao, Yotam Nitzan 等ICLR 2026 · 被引用 14 次
- UP2You: Fast Reconstruction of Yourself from Unconstrained Photo CollectionsZeyu Cai, Ziyang Li, Xiaoben Li, Boqian Li 等ICLR 2026 · 被引用 11 次
- IC-Custom: Diverse Image Customization via In-Context LearningYaowei Li, Xiaoyu Li, Zhaoyang Zhang, Yuxuan Bian 等ICLR 2026 · 被引用 10 次
- Dimension-Reduction Attack! Video Generative Models are Experts on Controllable Image SynthesisHengyuan Cao, Yutong Feng, Biao Gong, Yijing Tian 等NeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper50
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
相关 Paper
- CustomNet: Object Customization with Variable-Viewpoints in Text-to-Image Diffusion ModelsZiyang Yuan, Mingdeng Cao, Xintao Wang, Zhongang Qi 等ACM MM 2024 · 被引用 10 次
- MC^2: Multi-concept Guidance for Customized Multi-concept GenerationJiaxiu Jiang, Yabo Zhang, Kailai Feng, Xiaohe Wu 等CVPR 2025
- Encoder-based Domain Tuning for Fast Personalization of Text-to-Image ModelsRinon Gal, Moab Arar, Yuval Atzmon, Amit H. Bermano 等SIGGRAPH 2023 · 被引用 154 次
- SPAD: Spatially Aware Multi-View DiffusersYash Kant, Aliaksandr Siarohin, Ziyi Wu, Michael Vasilkovsky 等CVPR 2024
- Visual Persona: Foundation Model for Full-Body Human CustomizationJisu Nam, Soowon Son, Zhan Xu, Jing Shi 等CVPR 2025
