Multitwine: Multi-Object Compositing with Text and Layout Control
Gemma Canet Tarrés, Zhe Lin, Zhifei Zhang, He Zhang, Andrew Gilbert, John P. Collomosse, Soo Ye Kim
Abstract
We introduce the first generative model capable of simultaneous multi-object compositing, guided by both text and layout. Our model allows for the addition of multiple objects within a scene, capturing a range of interactions from simple positional relations (e.g., next to, in front of) to complex actions requiring reposing (e.g., hugging, playing guitar). When an interaction implies additional props, like 'taking a selfie', our model autonomously generates these * : corresponding authors. This work was a joint collaboration between Adobe and the University of Surrey, conducted during an internship of the main author at Adobe. It was partially supported by DECaDE under EPSRC Grant EP/T022485/1. supporting objects. By jointly training for compositing and subject-driven generation, also known as customization, we achieve a more balanced integration of textual and visual inputs for text-driven object compositing. As a result, we obtain a versatile model with state-of-the-art performance in both tasks. We further present a data generation pipeline leveraging visual and language models to effortlessly synthesize multimodal, aligned training data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Does FLUX Already Know How to Perform Physically Plausible Image Composition?Shilin Lu, Zhuming Lian, Zihan Zhou, Shaocong Zhang et al.ICLR 2026 · 34 citations
- SIGMA-Gen: Structure and Identity Guided Multi-Subject Assembly for Image GenerationOindrila Saha, Vojtech Krs, Radomir Mech, Subhransu Maji et al.ICLR 2026 · 5 citations
- PLACID: Identity-Preserving Multi-Object Compositing via Video Diffusion with Synthetic TrajectoriesGemma Canet Tarrés, Manel Baradad, Francesc Moreno-Noguer, Yumeng LiCVPR 2026 · 1 citation
- HOComp: Interaction-Aware Human-Object CompositionDong Liang, Jinyuan Jia, Yuhao Liu, Rynson W. H. LauNeurIPS 2025 · 1 citation
- BFS: Back-to-Front Layered Image Synthesis via Knowledge TransferKyoungkook Kang, Gyujin Sim, Sunghyun ChoSIGGRAPH 2026
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Text2Interact: High-Fidelity and Diverse Text-to-Two-Person Interaction GenerationQingxuan Wu, Zhiyang Dou, chuan guo, Yiming Huang et al.ICLR 2026 · 10 citations
- Customizable Image Synthesis with Multiple SubjectsZhiheng Liu, Yifei Zhang, Yujun Shen, Kecheng Zheng et al.NeurIPS 2023 · 20 citations
- Interact-Custom: Customized Human Object Interaction Image GenerationZhu Xu, Zhaowen Wang, Yuxin Peng, Yang LiuACM MM 2025 · 1 citation
- HECTOR: Hybrid Editable Compositional Object References for Video GenerationGuofeng Zhang, Angtian Wang, Jacob Fang, Liming Jiang et al.ICML 2026
- CompGS: Unleashing 2D Compositionality for Compositional Text-to-3D via Dynamically Optimizing 3D GaussiansChongjian Ge, Chenfeng Xu, Yuanfeng Ji, Chensheng Peng et al.CVPR 2025
