GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization
Chenxi Liu, Selena Ling, Alec Jacobson
摘要
Fig. 1. Given a prompt, default image generation may appear reasonable yet fail to match a user's creative intent (left). GimmBO enables users to explore weighted combinations of customization adapters (30 in this example) through preference feedback, producing an image that better reflects the desired style (right, corresponding adapter coefficients below). Readers are encouraged to zoom in on images for finer details throughout the paper.
Fine-tuning-based adaptation is widely used to customize diffusion-based image generation, leading to large collections of community-created adapters that capture diverse subjects and styles. Adapters derived from the same base model can be merged linearly, enabling the synthesis of new visual results within a vast and continuous design space. To explore this space, current workflows rely on manual slider-based tuning, an approach that scales poorly and makes merging coefficient selection difficult, even when the candidate set is limited to 20-30 adapters. We propose GimmBO to support interactive exploration of adapter merging for image generation through Preferential Bayesian Optimization (PBO). Motivated by observations from real-world usage, including sparsity and constrained coefficient ranges, we introduce a two-stage BO backend that improves sampling efficiency and convergence in high-dimensional spaces. We evaluate our approach with simulated users and a user study, demonstrating improved convergence, high success rates, and consistent gains over BO and line-search baselines, and further show the flexibility of the framework through several extensions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper33
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsChenlin Meng, Yutong He, Yang Song, Jiaming Song 等ICLR 2022 · 被引用 2,128 次
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel 等NeurIPS 2023 · 被引用 999 次
相关 Paper
- Personalized Image Generation via Human-in-the-loop Bayesian OptimizationRajalaxmi Rajagopalan, Debottam Dutta, Yu-Lin Wei, Romit Roy ChoudhuryICML 2026 · 被引用 2 次
- Stylus: Automatic Adapter Selection for Diffusion ModelsMichael Luo, Justin Wong, Brandon Trabucco, Yanping Huang 等NeurIPS 2024 · 被引用 27 次
- Efficient Visual Appearance Optimization by Learning from Prior PreferencesZhipeng Li, Yi-Chi Liao, Christian HolzUIST 2025
- Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion modelsKyungmin Lee, Sangkyung Kwak, Kihyuk Sohn, Jinwoo ShinNeurIPS 2024 · 被引用 13 次
- LoRAverse: A Submodular Framework to Retrieve Diverse Adapters for Diffusion ModelsMert Sonmezer, Matthew Zheng, Pinar YanardagICCV 2025 · 被引用 1 次
