Bifröst: 3D-Aware Image Compositing with Language Instructions
Lingxiao Li, Kaixiong Gong, Wei-Hong Li, Xili Dai, Tao Chen, Xiaojun Yuan, Xiangyu Yue
Abstract
This paper introduces Bifröst, a novel 3D-aware framework that is built upon diffusion models to perform instruction-based image composition. Previous methods concentrate on image compositing at the 2D level, which fall short in handling complex spatial relationships (, occlusion). Bifröst addresses these issues by training MLLM as a 2.5D location predictor and integrating depth maps as an extra condition during the generation process to bridge the gap between 2D and 3D, which enhances spatial comprehension and supports sophisticated spatial interactions. Our method begins by fine-tuning MLLM with a custom counterfactual dataset to predict 2.5D object locations in complex backgrounds from language instructions. Then, the image-compositing model is uniquely designed to process multiple types of input features, enabling it to perform high-fidelity image compositions that consider occlusion, depth blur, and image harmonization. Extensive qualitative and quantitative evaluations demonstrate that Bifröst significantly outperforms existing methods, providing a robust solution for generating realistically composited images in scenarios demanding intricate spatial understanding. This work not only pushes the boundaries of generative image compositing but also reduces reliance on expensive annotated datasets by effectively utilizing existing resources in innovative ways.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 466df365-932a-4bab-8a63-c1feb0a0be03Builds on32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP LatentsHan Lin, Jaemin Cho, Amir Zadeh, Chuan Li et al.NeurIPS 2025 · 9 citations
- 3D-aware Image Generation using 2D Diffusion ModelsJianfeng Xiang, Jiaolong Yang, Binbin Huang, Xin TongICCV 2023 · 82 citations
- BFS: Back-to-Front Layered Image Synthesis via Knowledge TransferKyoungkook Kang, Gyujin Sim, Sunghyun ChoSIGGRAPH 2026
- DreamComposer: Controllable 3D Object Generation via Multi-View ConditionsYunhan Yang, Yukun Huang, Xiaoyang Wu, Yuan-Chen Guo et al.CVPR 2024 · 3 citations
- ViHOI: Human-Object Interaction Synthesis with Visual PriorsSongjin Cai, Linjie Zhong, Ling Guo, Changxing DingCVPR 2026 · 2 citations
