CTRL-D: Controllable Dynamic 3D Scene Editing with Personalized 2D Diffusion
Kai He, Chin-Hsuan Wu, Igor Gilitschenski
Abstract
Input Scenes "Turn him into a joker" "Give him a hat" "Give him a pair of sunglasses" "Turn his body into a bronze statue" "Turn him into a robot" "Put him in a suit" "What if the man was painted in style?" "What if it was painted in style?" "Turn him into a Superman" Figure 1 . We present CTRL-D, a dynamic 3D scene editing framework that enables controllable, high-quality, consistent scene edits by editing only a single image using any 2D editing approach. Our framework is also compatible with both monocular and multi-camera scenes. Please see our project page for more dynamic visualizations: https://IHe-KaiI.github.io/CTRL-D/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- DEGauss: Defending Against Malicious 3D Editing for Gaussian SplattingLingzhuang Meng, Mingwen Shao, Yuanjian Qiao, Xiang LvNeurIPS 2025 · 4 citations
- Dynamic-eDiTor: Training-Free Text-Driven 4D Scene Editing with Multimodal Diffusion TransformerDong In Lee, Hyungjun Doh, Seunggeun Chi, Runlin Duan et al.CVPR 2026 · 3 citations
- Catalyst4D: High-Fidelity 3D-to-4D Scene Editing via Dynamic PropagationShifeng Chen, Yihui Li, Jun Liao, Hongyu Yang et al.CVPR 2026
- VDFE: Difference-Aware 3D Scene Editing with Non-Intrusive Video Diffusion Priors for Multi-View Consistency and EfficiencyChao Zhang, Fang Liu, Shuo Li, Yang Liu et al.CVPR 2026
Builds on47
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Edit3D: Elevating 3D Scene Editing with Attention-Driven Multi-Turn InteractivityPeng Zhou, Dunbo Cai, Yujian Du, Runqing Zhang et al.ACM MM 2024 · 3 citations
- TINKER: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene OptimizationCanyu Zhao, Xiaoman Li, Tianjian Feng, Zhiyue Zhao et al.ICLR 2026 · 9 citations
- 3DitScene: Editing Any Scene via Language-guided Disentangled Gaussian SplattingQihang Zhang, Yinghao Xu, Chaoyang Wang, Hsin-Ying Lee et al.ICLR 2025
- Drag Your Gaussian: Effective Drag-Based Editing with Score Distillation for 3D Gaussian SplattingYansong Qu, Dian Chen, Xinyang Li, Xiaofan Li et al.SIGGRAPH 2025 · 14 citations
- ShapeUP: Scalable Image-Conditioned 3D EditingInbar Gat, Dana Cohen-Bar, Guy Levy, Elad Richardson et al.SIGGRAPH 2026
