Composing Parts for Expressive Object Generation
Harsh Rangwani, Aishwarya Agarwal, Kuldeep Kulkarni, R. Venkatesh Babu, Srikrishna Karanam
Abstract
a full body portrait photo of a person wearing a dress with long hair and glasses PartComposer: hair, glasses, dress A photo of a yellow warbler PartComposer: beak → a hens beak a photo of a white swan, 8k, full hd PartComposer: crown → a crown of a peacock A young woman sits at a table in a beautiful, lush garden, painting by claude monet PartComposer: dress → a dress in Ukiyo-e style Figure 1. Base Generation (left) Comparison with Methods for Generations with Parts Details (right). PartComposer allows the generation of object images with specified attributes (color, style etc.) of parts for the chosen object in the base text prompt. StableDiffusion and Rich-Text [18] methods with part details either ignore the part instructions or generate inconsistent ojects.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52f337d7-8a4b-489d-b917-0b4851a1ace9Builds on32
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Score Jacobian Chaining: Lifting Pretrained 2D Diffusion Models for 3D GenerationHaochen Wang, Xiaodan Du, Jiahao Li, Raymond A. Yeh et al.CVPR 2023
- FashionComposer: Compositional Fashion Image GenerationSihui Ji, Yiyang Wang, Xi Chen, Xiaogang Xu et al.SIGGRAPH 2025 · 3 citations
- CompGS: Unleashing 2D Compositionality for Compositional Text-to-3D via Dynamically Optimizing 3D GaussiansChongjian Ge, Chenfeng Xu, Yuanfeng Ji, Chensheng Peng et al.CVPR 2025
- Training-Free Structured Diffusion Guidance for Compositional Text-to-Image SynthesisWeixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani et al.ICLR 2023 · 70 citations
- VSC: Visual Search Compositional Text-to-Image Diffusion ModelDo Huu Dat, Nam Hyeon-Woo, Po Yuan Mao, Tae-Hyun OhICCV 2025 · 1 citation
