Fourier Token Merging: Understanding and Capitalizing Frequency Domain for Efficient Image Generation
Jiesong Liu, Xipeng Shen
Abstract
Image generation requires intensive computations and faces challenges due to long latency. Exploiting redundancy in the input images and intermediate representations throughout the neural network pipeline is an effective way to accelerate image generation. Token merging (ToMe) exploits similarities among input tokens by clustering them and merges similar tokens into one, thus significantly reducing the number of tokens that are fed into the transformer block. This work introduces Fourier Token Merging , a new method for understanding and capitalizing frequency domain for efficient image generation. By introducing frequency token merging, we find that transforming the token into the frequency domain representation for clustering can better exert the ability of clustering based on the underlying redundancy after de-correlation. Through analytical and empirical studies, we demonstrate the benefits of using Fourier clustering over the original time domain clustering. We experimented Fourier Token Merging on the stable diffusion model, and the results show up to 25% reduction in latency without impairing image quality. The code is available at https://github.com/Fred1031/Fourier-Token-Merging .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7b7132c3-8fa3-48b7-aa2a-bb73c1ffbd85Builds on16
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
Related papers
- ToMA: Token Merge with Attention for Diffusion ModelsWenbo Lu, Shaoyi Zheng, Yuxuan Xia, Shengjie WangICML 2025
- FreqTS: Frequency-Aware Token Selection for Accelerating Diffusion ModelsXinye Yang, Yuxin Yang, Haoran Pang, Aaron Xuxiang Tian et al.AAAI 2025 · 2 citations
- TF-ATM: Training-Free Adaptive Token MergingXin Zhang, Weiying Xie, Yunsong Li, Xiaoyu Chen et al.ACM MM 2025
- Attend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT AccelerationHaipeng Fang, Sheng Tang, Juan Cao, Enshuo Zhang et al.CVPR 2025
- Importance-Based Token Merging for Efficient Image and Video GenerationHaoyu Wu, Jingyi Xu, Hieu Le, Dimitris SamarasICCV 2025 · 3 citations
