FineStyle: Fine-grained Controllable Style Personalization for Text-to-image Models
Gong Zhang, Kihyuk Sohn, Meera Hahn, Humphrey Shi, Irfan Essa
Abstract
Few-shot fine-tuning of text-to-image (T2I) generation models enables people to create unique images in their own style using natural languages without requiring extensive prompt engineering. However, fine-tuning with only a handful, as little as one, of image-text paired data prevents fine-grained control of style attributes at generation. In this paper, we present FineStyle, a few-shot fine-tuning method that allows enhanced controllability for style personalized text-to-image generation. To overcome the lack of training data for fine-tuning, we propose a novel concept-oriented data scaling that amplifies the number of image-text pair, each of which focuses on different concepts (e.g., objects) in the style reference image. We also identify the benefit of parameter-efficient adapter tuning of key and value kernels of ∗ Work
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68aab0f9-bf6c-4064-b7d5-c2db1bb3b07bCited by top-tier papers8
- RelationAdapter: Learning and Transferring Visual Relation with Diffusion TransformersYan Gong, Yiren Song, Yicheng Li, Chenglin Li et al.NeurIPS 2025 · 30 citations
- ImageRAG: Dynamic Image Retrieval for Reference-Guided Image GenerationRotem Shalev-Arkushin, Rinon Gal, Amit Bermano, Ohad FriedICLR 2026 · 25 citations
- IMG: Calibrating Diffusion Models via Implicit Multimodal GuidanceJiayi Guo, Chuanhao Yan, Xingqian Xu, Yulin Wang et al.ICCV 2025 · 4 citations
- DuoLoRA: Cycle-Consistent and Rank-Disentangled Content-Style PersonalizationAniket Roy, Shubhankar Borse, Shreya Kadambi, Debasmit Das et al.ICCV 2025 · 1 citation
- T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive GenerationChieh-Yun Chen, Min Shi, Gong Zhang, Humphrey ShiICCV 2025 · 1 citation
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Specialist Diffusion: Plug-and-Play Sample-Efficient Fine-Tuning of Text-to-Image Diffusion Models to Learn Any Unseen StyleHaoming Lu, Hazarapet Tunanyan, Kai Wang, Shant Navasardyan et al.CVPR 2023
- DomainGallery: Few-shot Domain-driven Image Generation by Attribute-centric FinetuningYuxuan Duan, Yan Hong, Bo Zhang, Jun Lan et al.NeurIPS 2024 · 2 citations
- Concept-Aware LoRA for Domain-Aligned Segmentation Dataset GenerationMinho Park, Sunghyun Park, Jungsoo Lee, Hyojin Park et al.CVPR 2026 · 1 citation
- Few-shot Image Generation Using Discrete Content RepresentationYan Hong, Li Niu, Jianfu Zhang, Liqing ZhangACM MM 2022 · 11 citations
- InstantBooth: Personalized Text-to-Image Generation without Test-Time FinetuningJing Shi, Wei Xiong, Zhe Lin, Hyun Joon JungCVPR 2024 · 115 citations
