Mod-Adapter: Tuning-Free and Versatile Multi-concept Personalization via Modulation Adapter
Weizhi Zhong, Huan Yang, Zheng Liu, Huiguo He, Zijian He, Xuesong Niu, Di Zhang, Guanbin Li
Abstract
Personalized text-to-image generation aims to synthesize images of user-provided concepts in diverse contexts. Despite recent progress in multi-concept personalization, most are limited to object concepts and struggle to customize abstract concepts (e.g., pose, lighting). Some methods have begun exploring multi-concept personalization supporting abstract concepts, but they require test-time fine-tuning for each new concept, which is time-consuming and prone to overfitting on limited training images. In this work, we propose a novel tuning-free method for multi-concept personalization that can effectively customize both object and abstract concepts without test-time fine-tuning. Our method builds upon the modulation mechanism in pre-trained Diffusion Transformers (DiTs) model, leveraging the localized and semantically meaningful properties of the modulation space. Specifically, we propose a novel module, Mod-Adapter, to predict concept-specific modulation direction for the modulation process of concept-related text tokens. It introduces vision-language cross-attention for extracting concept visual features, and Mixture-of-Experts (MoE) layers that adaptively map the concept features into the modulation space. Furthermore, to mitigate the training difficulty caused by the large gap between the concept image space and the modulation space, we introduce a VLM-guided pre-training strategy that leverages the strong image understanding capabilities of vision-language models to provide semantic supervision signals. For a comprehensive comparison, we extend a standard benchmark by incorporating abstract concepts. Our method achieves state-of-the-art performance in multi-concept personalization, supported by quantitative, qualitative, and human evaluations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c21237c4-e181-4c22-ad6d-cba17e4614d4Cited by top-tier papers7
- MultiCrafter: High-Fidelity Multi-Subject Generation via Disentangled Attention and Identity-Aware Preference AlignmentTao Wu, Yibo Jiang, Yehao Lu, Zhizhong Wang et al.CVPR 2026 · 4 citations
- Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept PersonalizationTsai-Shien Chen, Aliaksandr Siarohin, Gordon Guocheng Qian, Kuan-Chieh Jackson Wang et al.CVPR 2026 · 4 citations
- ChArtist: Generating Pictorial Charts with Unified Spatial and Subject ControlShishi Xiao, Tongyu Zhou, David H. Laidlaw, Gromit Yeuk-Yin ChanCVPR 2026 · 2 citations
- Premier: Personalized Preference Modulation with Learnable User Embedding in Text-to-Image GenerationZihao Wang, Yuxiang Wei, Xinpeng Zhou, Tianyu Zhang et al.CVPR 2026 · 1 citation
- DreamShot: Personalized Storyboard Synthesis with Video Diffusion PriorJunjia Huang, Binbin Yang, Pengxiang Yan, Jiyang Liu et al.CVPR 2026 · 1 citation
Builds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- TokenVerse: Versatile Multi-concept Personalization in Token Modulation SpaceDaniel Garibi, Shahar Yadin, Roni Paiss, Omer Tov et al.SIGGRAPH 2025 · 12 citations
- UniVerse: A Unified Modulation Framework for Segmentation-Free, Disentangled Multi-Concept PersonalizationQuynh Phung, Sandesh Ghimire, Minsi Hu, Chung-Chi Tsai et al.CVPR 2026
- Customization Assistant for Text-to-image GenerationYufan Zhou, Ruiyi Zhang, Jiuxiang Gu, Tong SunCVPR 2024 · 10 citations
- Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image ModelsGihyun Kwon, Simon Jenni, Dingzeyu Li, Joon-Young Lee et al.CVPR 2024
- TARA: Token-Aware LoRA for Composable Personalization in Diffusion ModelsYuqi Peng, Lingtao Zheng, Yufeng Yang, Yi Huang et al.AAAI 2026 · 2 citations
