Fourier Token Merging: Understanding and Capitalizing Frequency Domain for Efficient Image Generation
Jiesong Liu, Xipeng Shen
摘要
Image generation requires intensive computations and faces challenges due to long latency. Exploiting redundancy in the input images and intermediate representations throughout the neural network pipeline is an effective way to accelerate image generation. Token merging (ToMe) exploits similarities among input tokens by clustering them and merges similar tokens into one, thus significantly reducing the number of tokens that are fed into the transformer block. This work introduces Fourier Token Merging , a new method for understanding and capitalizing frequency domain for efficient image generation. By introducing frequency token merging, we find that transforming the token into the frequency domain representation for clustering can better exert the ability of clustering based on the underlying redundancy after de-correlation. Through analytical and empirical studies, we demonstrate the benefits of using Fourier clustering over the original time domain clustering. We experimented Fourier Token Merging on the stable diffusion model, and the results show up to 25% reduction in latency without impairing image quality. The code is available at https://github.com/Fred1031/Fourier-Token-Merging .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
相关 Paper
- ToMA: Token Merge with Attention for Diffusion ModelsWenbo Lu, Shaoyi Zheng, Yuxuan Xia, Shengjie WangICML 2025
- FreqTS: Frequency-Aware Token Selection for Accelerating Diffusion ModelsXinye Yang, Yuxin Yang, Haoran Pang, Aaron Xuxiang Tian 等AAAI 2025 · 被引用 2 次
- TF-ATM: Training-Free Adaptive Token MergingXin Zhang, Weiying Xie, Yunsong Li, Xiaoyu Chen 等ACM MM 2025
- Attend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT AccelerationHaipeng Fang, Sheng Tang, Juan Cao, Enshuo Zhang 等CVPR 2025
- Importance-Based Token Merging for Efficient Image and Video GenerationHaoyu Wu, Jingyi Xu, Hieu Le, Dimitris SamarasICCV 2025 · 被引用 3 次
