GeoRK2: Geometry-Guided Runge–Kutta Integration for Diffusion Transformer Acceleration
Chaoqun Sun, Zongjing Fu, Powei Chang, Jinpeng Zhang, JianXiang Xiang, Yukang Gao, Chenyu Wang
摘要
Diffusion transformer models deliver state-of-the-art image synthesis quality but suffer from prohibitively slow iterative sampling. Fewer sampling steps accelerate inference but inevitably distort intermediate features and degrade visual fidelity, while offering little relief in computational cost. To address these limitations, we present GeoRK2, a training-free framework that bridges numerical analysis and information geometry. GeoRK2 couples second-order Runge–Kutta (RK2) integration with a curvature-aware geometric flow derived from the model's noise predictions, establishing provably stable feature evolution dynamics under manifold-aware integration. By leveraging an empirical feature covariance–induced metric estimated from gradient covariances to capture intrinsic feature geometry and applying parallel transport along the manifold connection, GeoRK2 constrains error propagation under large-step integration, ensuring both numerical stability and structural fidelity. As a fully plug-and-play method, GeoRK2 requires no retraining and is compatible with mainstream pretrained diffusion transformers. Comprehensive experiments on image generation and super-resolution tasks across representative diffusion backbones (e.g., DiT-XL, HunyuanVideo, and FLUX.1-dev) demonstrate that GeoRK2 achieves 4–5× faster inference than baseline frameworks (FORA, TaylorSeer) with only marginal perceptual differences (∆FID ≈ 0.81), confirming its effectiveness and generality. All implementation details and code are provided in the supplementary material.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- Adaptive Spectral Feature Forecasting for Diffusion Sampling AccelerationJiaqi Han, Juntong Shi, Puheng Li, Haotian Ye 等CVPR 2026 · 被引用 4 次
- STORK: Faster Diffusion and Flow Matching Sampling by Resolving both Stiffness and Structure-DependenceZheng Tan, Weizhen Wang, Andrea L. Bertozzi, Ernest K. RyuICLR 2026 · 被引用 4 次
- Forecast Then Calibrate: Feature Caching as ODE for Efficient Diffusion TransformersShikang Zheng, Liang Feng, Xinyu Wang, Qinming Zhou 等AAAI 2026 · 被引用 10 次
- Just-in-Time: Training-Free Spatial Acceleration for Diffusion TransformersWenhao Sun, Ji Li, Zhaoqiang LiuCVPR 2026 · 被引用 5 次
- LookFlow: Training-Free and Efficient High-Resolution Image Synthesis via Dynamic Lookahead Guidance FlowYuan Zhou, Yan Zhang, Jianlong Chang, Xin Gu 等AAAI 2026
