Geometric Image Editing via Effects-Sensitive In-Context Inpainting with Diffusion Transformers
Shuo Zhang, Wenzhuo Wu, Huayu Zhang, Jiarong Cheng, Xianghao Zang, Chao Ban, Hao Sun, Zhongjiang He, Tianwei Cao, Kongming Liang, Zhanyu Ma
Abstract
Recent advances in diffusion models have significantly improved image editing. However, challenges persist in handling geometric transformations, such as translation, rotation, and scaling, particularly in complex scenes. Existing approaches suffer from two main limitations: (1) difficulty in achieving accurate geometric editing of object translation, rotation, and scaling; (2) inadequate modeling of intricate lighting and shadow effects, leading to unrealistic results. To address these issues, we propose GeoEdit, a framework that leverages in-context generation through a diffusion transformer module, which integrates geometric transformations for precise object edits. Moreover, we introduce Effects-Sensitive Attention, which enhances the modeling of intricate lighting and shadow effects for improved realism. To further support training, we construct RS-Objects, a large-scale geometric editing dataset containing over 120,000 high-quality image pairs, enabling the model to learn precise geometric editing while generating realistic lighting and shadows. Extensive experiments on public benchmarks demonstrate that GeoEdit consistently outperforms state-of-the-art methods in terms of visual quality, geometric accuracy, and realism.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a69b1a7-cce5-4edd-905f-26ab29e48ff3Builds on36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- 3D-Fixup: Advancing Photo Editing with 3D PriorsYen-Chi Cheng, Krishna Kumar Singh, Jae Shin Yoon, Alexander G. Schwing et al.SIGGRAPH 2025 · 4 citations
- Free-Form Scene Editor: Enabling Multi-Round Object Manipulation Like in a 3D EngineXincheng Shuai, Zhenyuan Qin, Henghui Ding, Dacheng TaoAAAI 2026 · 2 citations
- OBJECT 3DIT: Language-guided 3D-aware Image EditingOscar Michel, Anand Bhattad, Eli VanderBilt, Ranjay Krishna et al.NeurIPS 2023 · 79 citations
- GeoDiffusion: A Training-Free Framework for Accurate 3D Geometric Conditioning in Image GenerationPhillip Mueller, Talip Uenlue, Sebastian Schmidt, Marcel Kollovieh et al.ICCV 2025 · 2 citations
- Insert Anything: Image Insertion via In-Context Editing in DiTWensong Song, Hong Jiang, Zongxing Yang, Zheqiao Cheng et al.AAAI 2026 · 1 citation
