Vinedresser3D: Towards Agentic Text-guided 3D Editing
Yankuan Chi, Xiang Li, Zixuan Huang, James M.
摘要
Text-guided 3D editing aims to modify existing 3D assets using natural-language instructions. Current methods struggle to jointly understand complex prompts, automatically localize edits in 3D, and preserve unedited content. We introduce Vinedresser3D, an agentic framework for high-quality text-guided 3D editing that operates directly in the latent space of a native 3D generative model. Given a 3D asset and an editing prompt, Vinedresser3D uses a multimodal large language model to infer rich descriptions of the original asset, identify the parts to edit and the edit type (addition, modification, deletion), and generate decomposed structural and appearance-level text guidance. The agent then selects an informative view and applies an image editing model to obtain visual guidance. Finally, an inversion-based rectified-flow inpainting pipeline with an interleaved sampling module performs editing in the 3D latent space, enforcing prompt alignment while maintaining 3D coherence and unedited regions. Experiments on diverse 3D edits demonstrate that Vinedresser3D outperforms prior baselines in both automatic metrics and human preference studies, while enabling precise, coherent, and mask-free 3D editing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsChenlin Meng, Yutong He, Yang Song, Jiaming Song 等ICLR 2022 · 被引用 2,128 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
相关 Paper
- Easy3E: Feed-Forward 3D Asset Editing via Rectified Voxel FlowShimin Hu, Yuanyi Wei, Fei Zha, Yudong Guo 等CVPR 2026 · 被引用 7 次
- PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-Aware MaskJeongho Kim, Hoiyeong Jin, Sunghyun Park, Jaegul ChooICCV 2025 · 被引用 6 次
- TIP-Editor: An Accurate 3D Editor Following Both Text-Prompts And Image-PromptsJingyu Zhuang, Di Kang, Yan-Pei Cao, Guanbin Li 等SIGGRAPH 2024 · 被引用 59 次
- GG-Editor: Locally Editing 3D Avatars with Multimodal Large Language Model GuidanceYunqiu Xu, Linchao Zhu, Yi YangACM MM 2024 · 被引用 9 次
- Edit3D: Elevating 3D Scene Editing with Attention-Driven Multi-Turn InteractivityPeng Zhou, Dunbo Cai, Yujian Du, Runqing Zhang 等ACM MM 2024 · 被引用 3 次
