FRAG: Frequency Adapting Group for Diffusion Video Editing
Sunjae Yoon, Gwanhyeong Koo, Geonwoo Kim, Chang D. Yoo
Abstract
In video editing, the hallmark of a quality edit lies in its consistent and unobtrusive adjustment. Modification, when integrated, must be smooth and subtle, preserving the natural flow and aligning seamlessly with the original vision. Therefore, our primary focus is on overcoming the current challenges in high quality edit to ensure that each edit enhances the final product without disrupting its intended essence. However, quality deterioration such as blurring and flickering is routinely observed in recent diffusion video editing systems. We confirm that this deterioration often stems from high-frequency leak: the diffusion model fails to accurately synthesize high-frequency components during denoising process. To this end, we devise Frequency Adapting Group (FRAG) which enhances the video quality in terms of consistency and fidelity by introducing a novel receptive field branch to preserve high-frequency components during the denoising process. FRAG is performed in a model-agnostic manner without additional training and validates the effectiveness on video editing benchmarks (i.e., TGVE, DAVIS).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- TPC: Test-time Procrustes Calibration for Diffusion-based Human Image AnimationSunjae Yoon, Gwanhyeong Koo, Younghwan Lee, Chang Dong YooNeurIPS 2024 · 16 citations
- Frequency-Aware Flow Matching for High-Quality Image GenerationSucheng Ren, Qihang Yu, Ju He, Xiaohui Shen et al.CVPR 2026 · 6 citations
- TARO: Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning for Synchronized Video-To-Audio SynthesisTri Ton, Ji Woo Hong, Chang D. YooICCV 2025 · 1 citation
- ITA-MDT: Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-OnJi Woo Hong, Tri Ton, Trung X. Pham, Gwanhyeong Koo et al.CVPR 2025
- Occlusion-Robust Stylization for Drawing-Based 3D AnimationSunjae Yoon, Gwanhyeong Koo, Younghwan Lee, Ji Woo Hong et al.ICCV 2025
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- AnchorSync: Global Consistency Optimization for Long Video EditingZichi Liu, Yinggui Wang, Tao Wei, Chao MaACM MM 2025
- UFO: Enhancing Diffusion-Based Video Generation with a Uniform Frame OrganizerDelong Liu, Zhaohui Hou, Mingjie Zhan, Shihao Han et al.AAAI 2025
- FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal AttentionYu Lu, Yuanzhi Liang, Linchao Zhu, Yi YangNeurIPS 2024 · 101 citations
- Aligning Global Semantics and Local Textures in Generative Video EnhancementZhikai Chen, Fuchen Long, Zhaofan Qiu, Ting Yao et al.ICCV 2025 · 1 citation
- VideoDirector: Precise Video Editing via Text-to-Video ModelsYukun Wang, Longguang Wang, Zhiyuan Ma, Qibin Hu et al.CVPR 2025
