Learning Continuous 3D Words for Text-to-Image Generation
Ta Ying Cheng, Matheus Gadelha, Thibault Groueix, Matthew Fisher, Radomír Mech, Andrew Markham, Niki Trigoni
Abstract
Current controls over diffusion models (e.g., through text or ControlNet) for image generation fall short in recognizing abstract, continuous attributes like illumination direction or non-rigid shape change. In this paper, we present an approach for allowing users of text-to-image models to have fine-grained control of several attributes in an image. We do this by engineering special sets of input tokens that can be transformed in a continuous manner -we call them Continuous 3D Words. These attributes can, for example, be represented as sliders and applied jointly with text prompts for fine-grained control over image generation. Given only a single mesh and a rendering engine, we show that our approach can be adopted to provide continuous user control over several 3D-aware attributes, including time-of-day illumination, bird wing orientation, dollyzoom effect, and object poses. Our method is capable of conditioning image creation with multiple Continuous 3D Words and text descriptions simultaneously while adding no overhead to the generative process. Project
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c00e28e-ba5b-4209-ab98-15698cc28850Cited by top-tier papers10
- CamEdit: Continuous Camera Parameter Control for Photorealistic Image EditingXinran Qin, Zhixin Wang, Fan Li, Haoyu Chen et al.NeurIPS 2025 · 17 citations
- Kontinuous Kontext: Continuous Strength Control for Instruction-based Image EditingRishubh Parihar, Or Patashnik, Daniil Ostashev, Venkatesh Babu Radhakrishnan et al.CVPR 2026 · 15 citations
- SceneDesigner: Controllable Multi-Object Image Generation with 9-DoF Pose ManipulationZhenyuan Qin, Xincheng Shuai, Henghui DingNeurIPS 2025 · 11 citations
- ORIGEN: Zero-Shot 3D Orientation Grounding in Text-to-Image GenerationYunhong Min, Daehyeon Choi, Kyeongmin Yeo, Jihyun Lee et al.NeurIPS 2025 · 10 citations
- SeeThrough3D: Occlusion Aware 3D Control in Text-to-Image GenerationVaibhav Agrawal, Rishubh Parihar, Pradhaan Bhat, Ravi Kiran Sarvadevabhatla et al.CVPR 2026 · 5 citations
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Generating compositional scenes via Text-to-image RGBA Instance GenerationAlessandro Fontanella, Petru-Daniel Tudosiu, Yongxin Yang, Shifeng Zhang et al.NeurIPS 2024 · 13 citations
- AttriCtrl: A Generalizable Framework for Controlling Semantic Attribute Intensity in Diffusion ModelsDie Chen, Zhongjie Duan, Zhiwen Li, Cen Chen et al.ICLR 2026
- Control3D: Towards Controllable Text-to-3D GenerationYang Chen, Yingwei Pan, Yehao Li, Ting Yao et al.ACM MM 2023 · 54 citations
- CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image GenerationZixin Zhu, Kevin Duarte, Mamshad Nayeem Rizve, Chengyuan Xu et al.ICCV 2025
- SliderEdit: Continuous Image Editing with Fine-Grained Instruction ControlArman Zarei, Samyadeep Basu, Mobina Pournemat, Sayan Nag et al.CVPR 2026 · 12 citations
