LOOSECONTROL: Lifting ControlNet for Generalized Depth Conditioning
Shariq Farooq Bhat, Niloy J. Mitra, Peter Wonka
Abstract
We present LooseControl to allow generalized depth conditioning for diffusion-based image generation. ControlNet, the SOTA for depth conditioned image generation, produces remarkable results but relies on having access to detailed depth maps for guidance. Creating such exact depth maps, in many scenarios, is challenging. This paper introduces a generalized version of depth conditioning that enables new content creation workflows. Specifically, we allow (C1) scene boundary control for loosely specifying scenes with only boundary conditions, and (C2) 3D box control for specifying the target objects’ layout locations rather than the objects’ exact shape and appearance. Using LooseControl, along with text guidance, users can create complex environments (e.g., rooms, street views, etc.) by specifying only scene boundaries and locations of primary objects. Further, we provide two editing mechanisms to refine the results: (E1) 3D box editing enables the user to refine images by changing, adding, or removing boxes while freezing the image style. This yields minimal changes apart from changes induced by the edited boxes. (E2) Attribute editing proposes possible editing directions to change one particular aspect of the scene, such as the overall object density or a particular object. Tests and comparisons with baselines demonstrate the generality of our method. We believe that LooseControl can become an important design tool for easily creating complex environments and be extended to other forms of guidance channels. The project page can be found at https://shariqfarooq123.github.io/loose-control/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e1e54fbc-d19f-487c-8551-e5dcd83aa976Cited by top-tier papers38
- Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion ModelsZiyi Wu, Yulia Rubanova, Rishabh Kabra, Drew A. Hudson et al.NeurIPS 2024 · 32 citations
- Block and Detail: Scaffolding Sketch-to-Image GenerationVishnu Sarukkai, Lu Yuan, Mia Tang, Maneesh Agrawala et al.UIST 2024 · 23 citations
- Coin3D: Controllable and Interactive 3D Assets Generation with Proxy-Guided ConditioningWenqi Dong, Bangbang Yang, Lin Ma, Xiao Liu et al.SIGGRAPH 2024 · 18 citations
- SpaceBlender: Creating Context-Rich Collaborative Spaces Through Generative 3D Scene BlendingNels Numan, Shwetha Rajaram, Balasaravanan Thoravi Kumaravel, Nicolai Marquardt et al.UIST 2024 · 17 citations
- CineMaster: A 3D-Aware and Controllable Framework for Cinematic Text-to-Video GenerationQinghe Wang, Yawen Luo, Xiaoyu Shi, Xu Jia et al.SIGGRAPH 2025 · 13 citations
Builds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Build-A-Scene: Interactive 3D Layout Control for Diffusion-Based Image GenerationAbdelrahman Eldesokey, Peter WonkaICLR 2025 · 1 citation
- Simplifying Control Mechanism in Text-to-Image Diffusion ModelsZhida Feng, Li Chen, Yuenan Sun, Jiaxiang Liu et al.AAAI 2025
- DynFusion: Rethinking Condition Fusion for Adaptive Multi-Conditional Text-to-Image GenerationZheng Fang, Lichuan Xiang, Xu Cai, Bing Wang et al.CVPR 2026
- FreeControl: Efficient, Training-Free Structural Control via One-Step Attention ExtractionJiang Lin, Xinyu Chen, Song Wu, Zhiqiu Zhang et al.NeurIPS 2025 · 3 citations
