HomeDiffusion: Zero-Shot Object Customization with Multi-View Representation Learning for Indoor Scenes
Guoqiu Li, Jin Song, Yiyun Fei
Abstract
Recently, zero-shot object customization generation methods have rapidly developed and shown tremendous potential for applications. For instance, in the e-commerce domain, consumers can observe the visual effect of furniture placed within their personal living spaces or clothes worn on their own bodies. Many existing approaches perform object customization generation based on diffusion models and extracted reference object features. However, the generated object significantly diverges from the original reference object in details such as patterns and curves. Particularly for asymmetrical reference objects, the absence of comprehensive multi-viewpoint information prevents the generation of object poses that harmonize with the background scene. To address these shortcomings, we have constructed a novel dataset comprising multi-angle images of furniture and indoor scenes. Based on diffusion models, we introduce HomeDiffusion, which can leverage multi-viewpoint images of the same reference object to accurately generate visually harmonious object poses within specified areas of the background scene. During the diffusion process, we further extract high-fidelity details of the reference object and perform cross-attention with the noise latents in the latent space, thereby ensuring the preservation of details in the customized object generation. Extensive qualitative and quantitative experiments demonstrate that our method achieves superior performance over other existing zero-shot as well as few-shot object customization approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5cc5deb2-2d7b-4685-88c7-60a9c4db2543Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- AnyDoor: Zero-shot Object-level Image CustomizationXi Chen, Lianghua Huang, Yu Liu, Yujun Shen et al.CVPR 2024
- CustAny: Customizing Anything from A Single ExampleLingjie Kong, Kai Wu, Chengming Xu, Xiaobin Hu et al.CVPR 2025
- LoMOE: Localized Multi-Object Editing via Multi-DiffusionGoirik Chakrabarty, Aditya Chandrasekar, Ramya Hebbalaguppe, Prathosh APACM MM 2024 · 4 citations
- MVCustom: Multi-View Customized Diffusion via Geometric Latent Rendering and CompletionMinjung Shin, Hyunin Cho, Sooyeon Go, Jin-Hwa Kim et al.ICLR 2026 · 3 citations
- Harmonizing Visual and Textual Embeddings for Zero-Shot Text-to-Image CustomizationYeji Song, Jimyeong Kim, Wonhark Park, Wonsik Shin et al.AAAI 2025 · 6 citations
