Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling
Junha Hyung, Kinam Kim, Susung Hong, Min-Jung Kim, Jaegul Choo
2025Year
14Top-tier citations
Abstract
A close-up shot of a butterfly landing on the nose of a woman, highlighting her smile and the details of the butterfly's wings." "A close-up of a woman's face with colored powder exploding around her, creating an abstract splash of vibrant hues." Figure 1. Visual comparison of video quality between CFG (top row) and our STG method (bottom row). Best viewed in Acrobat Reader; click on the images to watch the videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e79ce7bf-2be5-4792-90a7-a1cb9825eaccCited by top-tier papers14
- Stochastic Self-Guidance for Training-Free Enhancement of Diffusion ModelsChubin Chen, Jiashu Zhu, Xiaokun Feng, Nisha Huang et al.ICLR 2026 · 44 citations
- Emergent Temporal Correspondences from Video Diffusion TransformersJisu Nam, Soowon Son, Dahyun Chung, Jiyoung Kim et al.NeurIPS 2025 · 30 citations
- Taming Video Models for 3D and 4D Generation via Zero-Shot Camera ControlChenxi Song, Yanming Yang, Tong Zhao, Ruibo Li et al.CVPR 2026 · 17 citations
- Guiding a Diffusion Transformer with the Internal Dynamics of ItselfXingyu Zhou, Qifan Li, Xiaobin Hu, Hai Chen et al.CVPR 2026 · 13 citations
- Syncphony: Synchronized Audio-to-Video Generation with Diffusion TransformersJibin Song, Mingi Kwon, Jaeseok Jeong, Youngjung UhICLR 2026 · 6 citations
Builds on17
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
Related papers
- EIDT-V: Exploiting Intersections in Diffusion Trajectories for Model-Agnostic, Zero-Shot, Training-Free Text-to-Video GenerationDiljeet Jagpal, Xi Chen, Vinay P. NamboodiriCVPR 2025
- Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video ContentQiuheng Wang, Yukai Shi, Jiarong Ou, Rui Chen et al.CVPR 2025
- GLIGEN: Open-Set Grounded Text-to-Image GenerationYuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu et al.CVPR 2023
- Relational Context Learning for Human-Object Interaction DetectionSanghyun Kim, Deunsol Jung, Minsu ChoCVPR 2023
- Real-Time High-Resolution Background MattingShanchuan Lin, Andrey Ryabtsev, Soumyadip Sengupta, Brian L. Curless et al.CVPR 2021
