EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM
Zhuofan Zong, Dongzhi Jiang, Bingqi Ma, Guanglu Song, Hao Shao, Dazhong Shen, Yu Liu, Hongsheng Li
摘要
Abstract Significant achievements in personalization of diffusion models have been witnessed. Conventional tuning-free methods mostly encode multiple reference images by averaging their image embeddings as the injection condition, but such an image-independent operation cannot perform interaction among images to capture consistent visual elements within multiple references. Although the tuning-based Low-Rank Adaptation (LoRA) can effectively extract consistent elements within multiple images through the training process, it necessitates specific finetuning for each distinct image group. This paper introduces EasyRef, a novel plug-and-play adaptation method that enables diffusion models to be conditioned on multiple reference im-ages and the text prompt. To effectively exploit consistent visual elements within multiple images, we leverage the multi-image comprehension and instruction-following capabilities of the multimodal large language model (MLLM), prompting it to capture consistent visual elements based on the instruction. Besides, injecting the MLLM's representations into the diffusion process through adapters can easily generalize to unseen domains, mining the consistent visual elements within unseen data. To mitigate computational costs and enhance fine-grained detail preservation, we introduce an efficient reference aggregation strategy and a progressive training scheme. Finally, we introduce MRBench, a new multi-reference image generation benchmark. Experimental results demonstrate EasyRef surpasses both tuning-free methods like IP-Adapter and tuning-based
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoTDongzhi Jiang, Ziyu Guo, Renrui Zhang, Zhuofan Zong 等NeurIPS 2025 · 被引用 181 次
- MoVA: Adapting Mixture of Vision Experts to Multimodal ContextZhuofan Zong, Bingqi Ma, Dazhong Shen, Guanglu Song 等NeurIPS 2024 · 被引用 110 次
- Exploring the Role of Large Language Models in Prompt Encoding for Diffusion ModelsBingqi Ma, Zhuofan Zong, Guanglu Song, Hongsheng Li 等NeurIPS 2024 · 被引用 57 次
- DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous DrivingYang Zhou, Hao Shao, Letian Wang, Zhuofan Zong 等ICLR 2026 · 被引用 20 次
- Neighbor GRPO: Contrastive ODE Policy Optimization Aligns Flow ModelsDailan He, Guanlin Feng, Xingtong Ge, Yazhe Niu 等CVPR 2026 · 被引用 15 次
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- CRAFT-LoRA: Content-Style Personalization via Rank-Constrained Adaptation and Training-Free FusionYu Li, Yujun Cai, Chi ZhangCVPR 2026 · 被引用 2 次
- TARA: Token-Aware LoRA for Composable Personalization in Diffusion ModelsYuqi Peng, Lingtao Zheng, Yufeng Yang, Yi Huang 等AAAI 2026 · 被引用 2 次
- Multimodal Instruction Tuning with Conditional Mixture of LoRAYing Shen, Zhiyang Xu, Qifan Wang, Yu Cheng 等ACL 2024
- Mixture-of-Subspaces in Low-Rank AdaptationTaiqiang Wu, Jiahao Wang, Zhe Zhao, Ngai WongEMNLP 2024 · 被引用 14 次
- EasyGen: Easing Multimodal Generation with BiDiffuser and LLMsXiangyu Zhao, Bo Liu, Qijiong Liu, Guangyuan Shi 等ACL 2024
