Aligning Global Semantics and Local Textures in Generative Video Enhancement
Zhikai Chen, Fuchen Long, Zhaofan Qiu, Ting Yao, Wengang Zhou, Jiebo Luo, Tao Mei
Abstract
Recent advances in video generation have demonstrated the utility of powerful diffusion models. One important direction among them is to enhance the visual quality of the AI-synthesized videos for artistic creation. Nevertheless, solely relying on the knowledge embedded in the pretrained video diffusion models might limit the generalization ability of local details (e.g., texture). In this paper, we address this issue by exploring the visual cues from a high-quality (HQ) image reference to facilitate visual details generation in video enhancement. We present GenVE, a new recipe of generative video enhancement framework that pursues the semantic and texture alignment between HQ image reference and denoised video in diffusion. Technically, GenVE first leverages an image diffusion model to magnify a key frame of the input video to attain a semanticsaligned HQ image reference. Then, a video controller is integrated into 3D-UNet to capture patch-level texture of the image reference to enhance fine-grained details generation at the corresponding region of low-quality (LQ) video. Moreover, a series of conditioning augmentation strategies are implemented for effective model training and algorithm robustness. Extensive experiments conducted on the public YouHQ40 and VideoLQ, as well as self-built AIGC-Vid dataset, quantitatively and qualitatively demonstrate the efficacy of our GenVE over the state-of-the-art video enhancement approaches. Source code is available at https://github.com/HiDream-ai/GenVE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 46889699-9cdc-47cc-b971-0cbfe3a35ea4Cited by top-tier papers3
- Distillation Models are Good Samplers for Diffusion Reinforcement LearningZunxu Liu, Aiqiu Wu, Zhaofan Qiu, Yingwei Pan et al.ICML 2026
- PS-SR: Pseudo-Single-Step Video Super-Resolution via Speculative DiffusionAiqiu Wu, Zhaofan Qiu, Ting Yao, Tao MeiCVPR 2026
- In-Context Generation with Regional Constraints for Instructional Video EditingZhongwei Zhang, Fuchen Long, Wei Li, Zhaofan Qiu et al.ICML 2026
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- VIDM: Video Implicit Diffusion ModelsKangfu Mei, Vishal M. PatelAAAI 2023 · 107 citations
- MIMIC: Mask-Injected Manipulation Video Generation with Interaction ControlTianxiao Chen, Jintao Rong, Huajin Chen, Jingya Wang et al.ICLR 2026
- FreeEnhance: Tuning-Free Image Enhancement via Content-Consistent Noising-and-Denoising ProcessYang Luo, Yiheng Zhang, Zhaofan Qiu, Ting Yao et al.ACM MM 2024 · 4 citations
- GenRec: Unifying Video Generation and Recognition with Diffusion ModelsZejia Weng, Xitong Yang, Zhen Xing, Zuxuan Wu et al.NeurIPS 2024 · 19 citations
- Enhanced Motion-aware Latent Diffusion Models for Video Frame InterpolationZhilin Huang, Chujun Qin, Yifei Xing, Wenming YangACM MM 2025
