Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers
Kanghyun Baek, Jaihyun Lew, Chaehun Shin, Jungbeom Lee, Sungroh Yoon
Abstract
Multimodal Diffusion Transformers (MM-DiTs) have achieved remarkable progress in text-toimage generation, yet they frequently suffer from concept omission, where specified objects or attributes fail to emerge in the generated image. By performing linear probing on text tokens, we demonstrate that text embeddings can distinguish a characteristic 'omission signal' representing the absence of target concepts. Leveraging this insight, we propose Omission Signal Intervention (OSI), which amplifies the omission signal to actively catalyze the generation of missing concepts. Comprehensive experiments on FLUX.1-Dev and SD3.5-Medium demonstrate that OSI significantly alleviates concept omission even in extreme scenarios. The code is available at https: //github.com/KangHyun-dsail/OSI .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 702761eb-3c17-4c80-8fde-e7fc692206c5Cited by top-tier papers1
Ask how each one uses itBuilds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- Rare Text Semantics Were Always There in Your Diffusion TransformerSeil Kang, Woojung Han, Dayun Ju, Seong Jae HwangNeurIPS 2025 · 6 citations
- Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion TransformersYuxuan Yao, Yuxuan Chen, Hui Li, Kaihui Cheng et al.ICML 2026
- Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image SynthesisBingda Tang, Boyang Zheng, Sayak Paul, Saining XieCVPR 2025
- DOS: Directional Object Separation in Text Embeddings for Multi-Object Image GenerationDongnam Byun, Jungwon Park, Jungmin Ko, Changin Choi et al.AAAI 2026
- Circuit Mechanisms for Spatial Relation Generation in Diffusion TransformersBinxu Wang, Jingxuan Fan, Xu PanCVPR 2026 · 4 citations
