Stylus: Automatic Adapter Selection for Diffusion Models
Michael Luo, Justin Wong, Brandon Trabucco, Yanping Huang, Joseph E. Gonzalez, Zhifeng Chen, Ruslan Salakhutdinov, Ion Stoica
摘要
Beyond scaling base models with more data or parameters, fine-tuned adapters provide an alternative way to generate high fidelity, custom images at reduced costs. As such, adapters have been widely adopted by open-source communities, accumulating a database of over 100K adapters-most of which are highly customized with insufficient descriptions. This paper explores the problem of matching the prompt to a set of relevant adapters, built on recent work that highlight the performance gains of composing adapters. We introduce Stylus, which efficiently selects and automatically composes task-specific adapters based on a prompt's keywords. Stylus outlines a three-stage approach that first summarizes adapters with improved descriptions and embeddings, retrieves relevant adapters, and then further assembles adapters based on prompts' keywords by checking how well they fit the prompt. To evaluate Stylus, we developed StylusDocs, a curated dataset featuring 75K adapters with pre-computed adapter embeddings. In our evaluation on popular Stable Diffusion checkpoints, Stylus achieves greater CLIP-FID Pareto efficiency and is twice as preferred, with humans and multimodal models as evaluators, over the base model. See stylus-diffusion.github.io for more.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- LinEAS: End-to-end Learning of Activation Steering with a Distributional LossPau Rodríguez, Michal Klein, Eleonora Gualdoni, Valentino Maiorca 等NeurIPS 2025 · 被引用 15 次
- Preventing Shortcuts in Adapter Training via Providing the ShortcutsAnujraaj Goyal, Guocheng Qian, Huseyin Coskun, Aarush Gupta 等NeurIPS 2025 · 被引用 4 次
- Beyond Adapter Retrieval: Latent Geometry-Preserving Composition via Sparse Task ProjectionPengfei Jin, Peng Shu, Sifan Song, Sekeun Kim 等AAAI 2026 · 被引用 2 次
- Contrastive Test-Time Composition of Multiple LoRA Models for Image GenerationTuna Han Salih Meral, Enis Simsar, Federico Tombari, Pinar YanardagICCV 2025 · 被引用 1 次
- LoRAverse: A Submodular Framework to Retrieve Diverse Adapters for Diffusion ModelsMert Sonmezer, Matthew Zheng, Pinar YanardagICCV 2025 · 被引用 1 次
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
相关 Paper
- SUR-adapter: Enhancing Text-to-Image Pre-trained Diffusion Models with Large Language ModelsShanshan Zhong, Zhongzhan Huang, Wushao Wen, Jinghui Qin 等ACM MM 2023 · 被引用 45 次
- Optimizing Prompts for Text-to-Image GenerationYaru Hao, Zewen Chi, Li Dong, Furu WeiNeurIPS 2023 · 被引用 303 次
- Diffusion in StyleMartin Nicolas Everaert, Marco Bocchio, Sami Arpa, Sabine Süsstrunk 等ICCV 2023 · 被引用 51 次
- DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative ModelsZijie J. Wang, Evan Montoya, David Munechika, Haoyang Yang 等ACL 2023 · 被引用 149 次
- Mv-Adapter: Multi-View Consistent Image Generation Made EasyZehuan Huang, Yuan-Chen Guo, Haoran Wang, Ran Yi 等ICCV 2025 · 被引用 14 次
