Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions
Stefan Andreas Baumann, Felix Krause, Michael Neumayr, Nick Stracke, Melvin Sevi, Vincent Tao Hu, Björn Ommer
2025Year
14Top-tier citations
Abstract
Figure 1. (a) We augment the prompt input of image generation models with fine-grained control of attribute expression in generated images (unmodified images are marked in green) in a subject-specific manner without additional cost during generation. (b, c) Previous methods only allow either fine-grained expression control or fine-grained localization when starting from the image generated from a basic prompt.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd23f65e-9f6a-4ff2-b825-1868ffdc4716Cited by top-tier papers14
- Interpreting the Weight Space of Customized Diffusion ModelsAmil Dravid, Yossi Gandelsman, Kuan-Chieh Wang, Rameen Abdal et al.NeurIPS 2024 · 40 citations
- Magnet: We Never Know How Text-to-Image Diffusion Models Work, Until We Learn How Vision-Language Models FunctionChenyi Zhuang, Ying Hu, Pan GaoNeurIPS 2024 · 28 citations
- Mod-Adapter: Tuning-Free and Versatile Multi-concept Personalization via Modulation AdapterWeizhi Zhong, Huan Yang, Zheng Liu, Huiguo He et al.ICLR 2026 · 17 citations
- IP-Composer: Semantic Composition of Visual ConceptsSara Dorfman, Dana Cohen-Bar, Rinon Gal, Daniel Cohen-OrSIGGRAPH 2025 · 4 citations
- Vibe Spaces for Creatively Connecting and Expressing Visual ConceptsHuzheng Yang, Katherine Xu, Andrew Lu, Michael D. Grossberg et al.CVPR 2026 · 4 citations
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Learning Continuous 3D Words for Text-to-Image GenerationTa Ying Cheng, Matheus Gadelha, Thibault Groueix, Matthew Fisher et al.CVPR 2024
- ITI-Gen: Inclusive Text-to-Image GenerationCheng Zhang, Xuanbai Chen, Siqi Chai, Chen Henry Wu et al.ICCV 2023 · 89 citations
- Devil is in the Detail: Towards Injecting Fine Details of Image Prompt in Image Generation via Conflict-free Guidance and Stratified AttentionKyungmin Jo, Jooyeol Yun, Jaegul ChooCVPR 2025
- Overcoming False Illusions in Real-World Face Restoration with Multi-Modal Guided Diffusion ModelKeda Tao, Jinjin Gu, Yulun Zhang, Xiucheng Wang et al.ICLR 2025
- Generating compositional scenes via Text-to-image RGBA Instance GenerationAlessandro Fontanella, Petru-Daniel Tudosiu, Yongxin Yang, Shifeng Zhang et al.NeurIPS 2024 · 13 citations
