Rare Text Semantics Were Always There in Your Diffusion Transformer
Seil Kang, Woojung Han, Dayun Ju, Seong Jae Hwang
摘要
Starting from flow- and diffusion-based transformers, Multi-modal Diffusion Transformers (MM-DiTs) have reshaped text-to-vision generation, gaining acclaim for exceptional visual fidelity. As these models advance, users continually push the boundary with imaginative or rare prompts, which advanced models still falter in generating, since their concepts are often too scarce to leave a strong imprint during pre-training. In this paper, we propose a simple yet effective intervention that surfaces rare semantics inside MM-DiTs without additional training steps, data, denoising-time optimization, or reliance on external modules (e.g., large language models). In particular, the joint-attention mechanism intrinsic to MM-DiT sequentially updates text embeddings alongside image embeddings throughout transformer blocks. We find that by mathematically expanding representational basins around text token embeddings via variance scale-up before the joint-attention blocks, rare semantics clearly emerge in MM-DiT's outputs. Furthermore, our results generalize effectively across text-to-vision tasks, including text-to-image, text-to-video, and text-driven image editing. Our work invites generative models to reveal the semantics that users intend, once hidden yet ready to surface.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases ThemWoojung Han, Seil Kang, Youngjun Jun, Min-Hung Chen 等ICML 2026 · 被引用 2 次
- Efficient Weighted Sampling via Score-based Generative ModelsHeasung Kim, Taekyun Lee, Hyeji Kim, Gustavo De VecianaCVPR 2026 · 被引用 1 次
- When Do Diffusion Models Learn to Generate Multiple Objects?Yujin Jeong, Arnas Uselis, Iro Laina, Seong Joon Oh 等ICML 2026
- RAIGen: Rare Attribute Identification in Text-to-Image Generative ModelsSilpa Vadakkeeveetil Sreelatha, Dan Wang, Serge Belongie, Muhammad Awais 等ICML 2026
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
相关 Paper
- Diagnosing and Correcting Concept Omission in Multimodal Diffusion TransformersKanghyun Baek, Jaihyun Lew, Chaehun Shin, Jungbeom Lee 等ICML 2026 · 被引用 1 次
- Seg4Diff: Unveiling Open-Vocabulary Semantic Segmentation in Text-to-Image Diffusion TransformersChaehyun Kim, Heeseong Shin, Eunbeen Hong, Heeji Yoon 等NeurIPS 2025 · 被引用 6 次
- QK-Edit: Revisiting Attention-based Injection in MM-DiT for Image and Video EditingTiancheng Shen, Zilong Huang, Xiangtai Li, Zhijie Lin 等ICCV 2025 · 被引用 2 次
- Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion TransformersYuxuan Yao, Yuxuan Chen, Hui Li, Kaihui Cheng 等ICML 2026
- Unified Safe In-context Image Generation in Multimodal Diffusion Transformers via Restricting Unsafe Information FlowsXiang Yang, Feifei Li, Mi Zhang, Geng Hong 等ICML 2026
