Instant3dit: Multiview Inpainting for Fast Editing of 3D Objects
Amir Barda, Matheus Gadelha, Vladimir G. Kim, Noam Aigerman, Amit H. Bermano, Thibault Groueix
Abstract
Original Mesh Input mesh + mask "An elven warrior" Original Mesh 25 sec. "Man wearing a medieval helmet" Mesh Adaptive Remeshing NeRF Mesh Adaptive Remeshing Figure 1. Our method takes as input a 3D object along with a 3D mask (first column) and a text prompt, and uses our multiview inpainting diffusion model to consistently paint the mask in four rendered views of the object. Off-the-shelf reconstructors can be used on the multiview output to give an NeRF, a Gaussian Splat (second column), or a mesh (third column) that can be used along with adaptive remeshing to ensure the unmasked region is exactly preserved e.g. topology, uvs, (fourth and fifth column). This feedforward approach is orders of magnitude faster than previous works in generative 3D editing, taking just ≈ 3 seconds per multiview edit, then 0.7 seconds to reconstruct a GS or a NeRF, 3 seconds for a mesh, and ≈ 20 seconds for optional mesh post-processing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7b9f8fa5-c6f4-43de-a4e0-214d1d7306fcCited by top-tier papers12
- Nano3D: A Training-Free Approach for Efficient 3D Editing Without MasksJunliang Ye, Shenghao Xie, Ruowen Zhao, Zhengyi Wang et al.ICLR 2026 · 32 citations
- SpaceControl: Introducing Test-Time Spatial Control to 3D Generative ModelingElisabetta Fedele, Francis Engelmann, Ian Huang, Or Litany et al.ICLR 2026 · 11 citations
- AnchorFlow: Training-Free 3D Editing via Latent Anchor-Aligned FlowsZhenglin Zhou, Fan Ma, Chengzhuo Gui, Xiaobo Xia et al.CVPR 2026 · 11 citations
- Towards Scalable and Consistent 3D EditingRuihao Xia, Yang Tang, Pan ZhouICML 2026 · 8 citations
- Easy3E: Feed-Forward 3D Asset Editing via Rectified Voxel FlowShimin Hu, Yuanyi Wei, Fei Zha, Yudong Guo et al.CVPR 2026 · 7 citations
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- DreamCatalyst: Fast and High-Quality 3D Editing via Controlling Editability and Identity PreservationJiwook Kim, Seonho Lee, Jaeyo Shin, Jiho Choi et al.ICLR 2025
- ATT3D: Amortized Text-to-3D Object SynthesisJonathan Lorraine, Kevin Xie, Xiaohui Zeng, Chen-Hsuan Lin et al.ICCV 2023 · 100 citations
- GaussianEditor: Editing 3D Gaussians Delicately with Text InstructionsJunjie Wang, Jiemin Fang, Xiaopeng Zhang, Lingxi Xie et al.CVPR 2024 · 65 citations
- EpiDiff: Enhancing Multi-View Synthesis via Localized Epipolar-Constrained DiffusionZehuan Huang, Hao Wen, Junting Dong, Yaohui Wang et al.CVPR 2024
- Edit3D: Elevating 3D Scene Editing with Attention-Driven Multi-Turn InteractivityPeng Zhou, Dunbo Cai, Yujian Du, Runqing Zhang et al.ACM MM 2024 · 3 citations
