Insert Anything: Image Insertion via In-Context Editing in DiT
Wensong Song, Hong Jiang, Zongxing Yang, Zheqiao Cheng, Ruijie Quan, Yi Yang
摘要
This work presents Insert Anything, a unified framework for reference-based image insertion that seamlessly integrates objects from reference images into target scenes under flexible, user-specified control guidance. Instead of training separate models for individual tasks, our approach is trained once on our new AnyInsertion dataset, the first open-source large-scale dataset specifically designed for reference image–based image editing, comprising 136K prompt-image pairs covering diverse tasks such as person, object, and garment insertion--and effortlessly generalizes to a wide range of insertion scenarios. Such a challenging setting requires capturing both identity features and fine-grained details, while allowing versatile local adaptations in style, color, and texture. To this end, we propose to leverage the multimodal attention of the Diffusion Transformer (DiT) to support both mask- and text-guided editing. Furthermore, we introduce an in-context editing mechanism that treats the reference image as contextual information, employing two prompting strategies to harmonize the inserted elements with the target scene while faithfully preserving their distinctive features. Extensive experiments on AnyInsertion, DreamBooth, and VTON-HD benchmarks demonstrate that our method consistently outperforms existing alternatives, underscoring its great potential in real-world applications such as creative content generation, virtual try-on, and scene composition.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- Unified In-Context Video EditingZixuan Ye, Xuanhua He, Quande Liu, Qiulin Wang 等ICLR 2026 · 被引用 37 次
- Does FLUX Already Know How to Perform Physically Plausible Image Composition?Shilin Lu, Zhuming Lian, Zihan Zhou, Shaocong Zhang 等ICLR 2026 · 被引用 34 次
- RelationAdapter: Learning and Transferring Visual Relation with Diffusion TransformersYan Gong, Yiren Song, Yicheng Li, Chenglin Li 等NeurIPS 2025 · 被引用 30 次
- DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion TransformersZitong Wang, Hang Zhao, Qianyu Zhou, Xuequan Lu 等CVPR 2026 · 被引用 26 次
- ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance GenerationRuihang Xu, Dewei Zhou, Fan Ma, Yi YangICLR 2026 · 被引用 19 次
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- LightGlue: Local Feature Matching at Light SpeedPhilipp Lindenberger, Paul-Edouard Sarlin, Marc PollefeysICCV 2023 · 被引用 936 次
相关 Paper
- RAA: Achieving Interactive Remove/Add Anything via Fully Synthetic DataDelong Liu, Haotian Hou, Zhaohui Hou, Shihao Han 等AAAI 2026
- Teleportraits: Training-Free People Insertion Into Any SceneJialu Gao, K. J. Joseph, Fernando De la TorreICCV 2025
- Personalize Anything for Free with Diffusion TransformerHaoran Feng, Zehuan Huang, Lin Li, Lu ShengAAAI 2026 · 被引用 1 次
- Inpaint-Anywhere: Zero-Shot Multi-Identity Inpainting with Efficient Diffusion TransformerJunsheng Luan, Lei Zhao, Wei XingAAAI 2026
- DreamFuse: Adaptive Image Fusion with Diffusion TransformerJunjia Huang, Pengxiang Yan, Jiyang Liu, Jie Wu 等ICCV 2025 · 被引用 3 次
