Teleportraits: Training-Free People Insertion Into Any Scene
Jialu Gao, K. J. Joseph, Fernando De la Torre
Abstract
The task of realistically inserting a human from a reference image into a background scene is highly challenging, requiring the model to (1) determine the correct location and poses of the person and (2) perform high-quality personalization conditioned on the background. Previous approaches often treat them as separate problems, overlooking their interconnections, and typically rely on training to achieve high performance. In this work, we introduce a unified training-free pipeline that leverages pre-trained text-to-image diffusion models. We show that diffusion models inherently possess the knowledge to place people in complex scenes without requiring task-specific training. By combining inversion techniques with classifier-free guidance, our method achieves affordance-aware global editing, seamlessly inserting people into scenes. Furthermore, our proposed mask-guided self-attention mechanism ensures high-quality personalization, preserving the subject's identity, clothing, and body features from just a single reference image. To the best of our knowledge, we are the first to perform realistic human insertions into scenes in a training-free manner and achieve state-of-the-art results in diverse composite scene images with excellent identity preservation in backgrounds and subjects.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15c5a66f-d39d-4fd8-8f20-67d678c532baBuilds on36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Energy-Guided Optimization for Personalized Image Editing with Pretrained Text-to-Image Diffusion ModelsRui Jiang, Xinghe Fu, Guangcong Zheng, Teng Li et al.AAAI 2025 · 2 citations
- Putting People in Their Place: Affordance-Aware Human Insertion into ScenesSumith Kulal, Tim Brooks, Alex Aiken, Jiajun Wu et al.CVPR 2023
- Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion ModelsYoad Tewel, Rinon Gal, Dvir Samuel, Yuval Atzmon et al.ICLR 2025
- FreeInsert: Personalized Object Insertion with Geometric and Style ControlYuhong Zhang, Han Wang, Yiwen Wang, Rong Xie et al.ACM MM 2025 · 1 citation
- Insert Anything: Image Insertion via In-Context Editing in DiTWensong Song, Hong Jiang, Zongxing Yang, Zheqiao Cheng et al.AAAI 2026 · 1 citation
