LaTexBlend: Scaling Multi-concept Customized Generation with Latent Textual Blending
Jian Jin, Zhenbo Yu, Yang Shen, Zhenyong Fu, Jian Yang
Abstract
tioned after the text encoder and a linear projection. LA-TEXBLEND customizes each concept individually, storing them in a concept bank with a compact representation of latent textual features that captures sufficient concept information to ensure high fidelity. At inference, concepts from the bank can be freely and seamlessly combined in the latent textual space, offering two key merits for multiconcept generation: 1) excellent scalability, and 2) significant reduction of denoising deviation, preserving coherent layouts. Extensive experiments demonstrate that LATEXBLEND can flexibly integrate multiple customized concepts with harmonious structures and high subject fidelity, substantially outperforming baselines in both generation quality and computational efficiency. Project page: https://jinjianrick.github.io/latexblend/ This CVPR paper is the Open Access version, provided by the Computer Vision Foundation. Except for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore. MuDI Concept bank OMG Mix-of-Show V 2 * bear plushie sitting on V 1 * chair. V 7 * dog playing V 8 * guitar, surrounded by V 6 * flower, with V 10 * lighthouse in the background. V 11 * cat sitting next to V 12 * teddybear, with V 6 * flower blooming beside them, with V 10 * lighthouse and V 9 * barn in the background. Two kids wearing V 3 * jacket and V 4 * shoes, playing with V 5 * dog.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b449df33-c50b-4425-bc3c-e1e152113f6bCited by top-tier papers4
- Mod-Adapter: Tuning-Free and Versatile Multi-concept Personalization via Modulation AdapterWeizhi Zhong, Huan Yang, Zheng Liu, Huiguo He et al.ICLR 2026 · 17 citations
- Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven GenerationAbdelrahman Eldesokey, Aleksandar Cvejic, Bernard Ghanem, Peter WonkaNeurIPS 2025 · 6 citations
- CaricHarmony: Contrastive Diffusion Paths for Identity-Preserving Caricature SynthesisDongyu Wang, Dar-Yen Chen, Yi-Zhe SongCVPR 2026
- DVAR: Dynamic Visual Autoregressive Modeling for Image Super-ResolutionYu Zheng, Kai Zhang, Wei Zhu, Qingguo Liu et al.CVPR 2026
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- MultiBooth: Towards Generating All Your Concepts in an Image from TextChenyang Zhu, Kai Li, Yue Ma, Chunming He et al.AAAI 2025 · 52 citations
- MC^2: Multi-concept Guidance for Customized Multi-concept GenerationJiaxiu Jiang, Yabo Zhang, Kailai Feng, Xiaohe Wu et al.CVPR 2025
- Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image ModelsGihyun Kwon, Simon Jenni, Dingzeyu Li, Joon-Young Lee et al.CVPR 2024
- Concept Conductor: Orchestrating Multiple Personalized Concepts in Text-to-Image SynthesisZebin Yao, Fangxiang Feng, Ruifan Li, Xiaojie WangAAAI 2025 · 3 citations
- TweedieMix: Improving Multi-Concept Fusion for Diffusion-based Image/Video GenerationGihyun Kwon, Jong Chul YeICLR 2025
