Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers
Chaofan Gan, Zicheng Zhao, Yuanpeng Tu, Xi Chen, Ziran Qin, Tieyuan Chen, Mehrtash Harandi, Weiyao Lin
Abstract
Massive Activations (MAs) are a well-documented phenomenon across Transformer architectures, and prior studies in both LLMs and ViTs have shown that they play a substantial role in shaping model behavior. However, the nature and function of MAs within Diffusion Transformers (DiTs) remain largely unexplored. In this work, we systematically investigate these activations to elucidate their role in visual generation. We found that these massive activations occur across all spatial tokens, and their distribution is modulated by the input timestep embeddings. Importantly, our investigations further demonstrate that these massive activations play a key role in local detail synthesis, while having minimal impact on the overall semantic content of output. Building on these insights, we propose Detail Guidance (DG), a MAs-driven, training-free self-guidance strategy to explicitly enhance local detail fidelity for DiTs. Specifically, DG constructs a degraded ``detail-deficient'' model by disrupting MAs and leverages it to guide the original network toward higher-quality detail synthesis. Our DG can seamlessly integrate with Classifier-Free Guidance (CFG), enabling joint enhancement of detail fidelity and prompt alignment. Extensive experiments demonstrate that our DG consistently improves local detail quality across various pre-trained DiTs (, SD3, SD3.5, and Flux).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Generalization of Diffusion Models Arises with a Balanced Representation SpaceZekai Zhang, Xiao Li, Xiang Li, Lianghe Shi et al.ICLR 2026 · 14 citations
- VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video UnderstandingZhihao He, Tieyuan Chen, Kangyu Wang, Ziran Qin et al.ICML 2026 · 3 citations
Related papers
- Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive ActivationsChaofan Gan, Yuanpeng Tu, Xi Chen, Tieyuan Chen et al.NeurIPS 2025 · 22 citations
- Improving Sample Quality of Diffusion Models Using Self-Attention GuidanceSusung Hong, Gyuseong Lee, Wooseok Jang, Seungryong KimICCV 2023 · 167 citations
- Guiding a Diffusion Model by Swapping Its TokensWeijia Zhang, Yuehao Liu, Shanyan Guan, Wu Ran et al.CVPR 2026 · 2 citations
- Guiding Diffusion Models with Semantically Degraded ConditionsShilong Han, Yuming Zhang, Hongxia WangCVPR 2026 · 1 citation
- Improving Diffusion Generalization with Weak-to-Strong Segmented GuidanceLiangyu Yuan, Yufei Huang, Mingkun Lei, Tong Zhao et al.CVPR 2026 · 1 citation
