Dual Optimal Transport for Multi-Concept Composition: Structural Alignment and Texture Injection in Diffusion Models
Hao Fu, Tianyu Su, Meng Liu, Chenfang Yang, Tian Gan
Abstract
Diffusion models have shown impressive capabilities in text-to-image synthesis. However, multi-concept personalized generation remains challenging, particularly in aligning multiple reference concepts while preserving fidelity. To address this, we propose a novel Sketch-to-Rendering framework that leverages for structure alignment and texture injection. Our approach consists of two key components: , which ensures shape alignment by using mass-preserving OT for spatial consistency, and , which leverages low-frequency structure alignment to inject high-frequency texture details via OT-based residual transfer, thereby preserving texture fidelity without distorting structure. Extensive experiments demonstrate that our method significantly enhances both conceptual fidelity and visual quality. Ablation studies further validate the effectiveness of our optimal transport guidance and the decoupling of structure and texture during the generation process. Our code is available at https://github.com/fuhao7i/OTComp.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b21d7a9f-50d9-46c0-a8d2-859f4ade1ea6Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- SAM 3: Segment Anything with ConceptsNicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath et al.ICLR 2026 · 1,103 citations
Related papers
- Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image ModelsGihyun Kwon, Simon Jenni, Dingzeyu Li, Joon-Young Lee et al.CVPR 2024
- Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion modelsKyungmin Lee, Sangkyung Kwak, Kihyuk Sohn, Jinwoo ShinNeurIPS 2024 · 13 citations
- OmniPortrait: Fine-Grained Personalized Portrait Synthesis via Pivotal OptimizationDongxu Yue, Bo Lin, Yao Tang, Jiajun Liang et al.ICLR 2026
- Rethinking Diffusion Bridge Model with Dual Alignments for Medical Image SynthesisJinbao Wei, Yuhang Chen, Zhijie Wang, Gang Yang et al.ACM MM 2025 · 3 citations
- UniversalBooth: Model-Agnostic Personalized Text-To-Image GenerationSonghua Liu, Ruonan Yu, Xinchao WangICCV 2025 · 2 citations
