OBJECT 3DIT: Language-guided 3D-aware Image Editing
Oscar Michel, Anand Bhattad, Eli VanderBilt, Ranjay Krishna, Aniruddha Kembhavi, Tanmay Gupta
摘要
Existing image editing tools, while powerful, typically disregard the underlying 3D geometry from which the image is projected. As a result, edits made using these tools may become detached from the geometry and lighting conditions that are at the foundation of the image formation process. In this work, we formulate the newt ask of language-guided 3D-aware editing, where objects in an image should be edited according to a language instruction in context of the underlying 3D scene. To promote progress towards this goal, we release OBJECT: a dataset consisting of 400K editing examples created from procedurally generated 3D scenes. Each example consists of an input image, editing instruction in language, and the edited image. We also introduce 3DIT : single and multi-task models for four editing tasks. Our models show impressive abilities to understand the 3D composition of entire scenes, factoring in surrounding objects, surfaces, lighting conditions, shadows, and physically-plausible object configurations. Surprisingly, training on only synthetic scenes from OBJECT, editing capabilities of 3DIT generalize to real-world images.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D PriorsChenyang Ma, Kai Lu, Ta Ying Cheng, Niki Trigoni 等NeurIPS 2024 · 被引用 82 次
- Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion ModelsZiyi Wu, Yulia Rubanova, Rishabh Kabra, Drew A. Hudson 等NeurIPS 2024 · 被引用 32 次
- Shadows Don't Lie and Lines Can't Bend! Generative Models Don't know Projective Geometry...for NowAyush Sarkar, Hanlin Mai, Amitabh Mahapatra, Svetlana Lazebnik 等CVPR 2024 · 被引用 22 次
- LuxDiT: Lighting Estimation with Video Diffusion TransformerRuofan Liang, Kai He, Zan Gojcic, Igor Gilitschenski 等NeurIPS 2025 · 被引用 20 次
- Diffusion Handles Enabling 3D Edits for Diffusion Models by Lifting Activations to 3DKarran Pandey, Paul Guerrero, Matheus Gadelha, Yannick Hold-Geoffroy 等CVPR 2024 · 被引用 19 次
它引用的顶会 Paper40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Free-Form Scene Editor: Enabling Multi-Round Object Manipulation Like in a 3D EngineXincheng Shuai, Zhenyuan Qin, Henghui Ding, Dacheng TaoAAAI 2026 · 被引用 2 次
- ShapeTalk: A Language Dataset and Framework for 3D Shape Edits and DeformationsPanos Achlioptas, Ian Huang, Minhyuk Sung, Sergey Tulyakov 等CVPR 2023
- Geometric Image Editing via Effects-Sensitive In-Context Inpainting with Diffusion TransformersShuo Zhang, Wenzhuo Wu, Huayu Zhang, Jiarong Cheng 等ICLR 2026 · 被引用 1 次
- 3DitScene: Editing Any Scene via Language-guided Disentangled Gaussian SplattingQihang Zhang, Yinghao Xu, Chaoyang Wang, Hsin-Ying Lee 等ICLR 2025
- Insert Anything: Image Insertion via In-Context Editing in DiTWensong Song, Hong Jiang, Zongxing Yang, Zheqiao Cheng 等AAAI 2026 · 被引用 1 次
