Contrastive Test-Time Composition of Multiple LoRA Models for Image Generation
Tuna Han Salih Meral, Enis Simsar, Federico Tombari, Pinar Yanardag
摘要
Low-Rank Adaptation (LoRA) has emerged as a powerful and popular technique for personalization, enabling efficient adaptation of pre-trained image generation models for specific tasks without comprehensive retraining. While employing individual pre-trained LoRA models excels at representing single concepts, such as those representing a specific dog or a cat, utilizing multiple LoRA models to capture a variety of concepts in a single image still poses a significant challenge. Existing methods often fall short, primarily because the attention mechanisms within different LoRA models overlap, leading to scenarios where one concept may be completely ignored (e.g., omitting the dog) or where concepts are incorrectly combined (e.g., producing an image of two cats instead of one cat and one dog). We introduce CLoRA, a training-free approach that addresses these limitations by updating the attention maps of multiple LoRA models at test-time, and leveraging the attention maps to create semantic masks for fusing latent representations. This enables the generation of composite images that accurately reflect the characteristics of each LoRA. Our comprehensive qualitative and quantitative evaluations demonstrate that CLoRA significantly outperforms existing methods in multi-concept image generation using LoRAs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SplitFlux: Learning to Decouple Content and Style from a Single ImageYitong Yang, Yinglin Wang, Changshuo Wang, Yongjun Zhang 等CVPR 2026 · 被引用 5 次
- CRAFT-LoRA: Content-Style Personalization via Rank-Constrained Adaptation and Training-Free FusionYu Li, Yujun Cai, Chi ZhangCVPR 2026 · 被引用 2 次
- Continual Personalization for Diffusion ModelsYu-Chien Liao, Jr-Jen Chen, Chi-Pin Huang, Ci-Siang Lin 等ICCV 2025 · 被引用 2 次
- Compression as Adaptation: Implicit Visual Representation with Diffusion Foundation ModelsZongyu Guo, Jiajun He, Zhaoyang Jia, Xiaoyi Zhang 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- TARA: Token-Aware LoRA for Composable Personalization in Diffusion ModelsYuqi Peng, Lingtao Zheng, Yufeng Yang, Yi Huang 等AAAI 2026 · 被引用 2 次
- Cached Multi-Lora Composition for Multi-Concept Image GenerationXiandong Zou, Mingzhu Shen, Christos-Savvas Bouganis, Yiren ZhaoICLR 2025
- UnZipLoRA: Separating Content and Style from a Single ImageChang Liu, Viraj Shah, Aiyu Cui, Svetlana LazebnikICCV 2025 · 被引用 31 次
- LoRAverse: A Submodular Framework to Retrieve Diverse Adapters for Diffusion ModelsMert Sonmezer, Matthew Zheng, Pinar YanardagICCV 2025 · 被引用 1 次
- qa-FLoRA: Data-free query-adaptive Fusion of LoRAs for LLMsShreya Shukla, Aditya Sriram, Milinda Kuppur Narayanaswamy, Hiteshi JainAAAI 2026
