Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion Model
Yang Yang, Siming Zheng, Qirui Yang, Jinwei Chen, Boxi Wu, Xiaofei He, Deng Cai, Bo Li, Peng-Tao Jiang
Abstract
Diffusion models have recently emerged as powerful tools for camera simulation, enabling both geometric transformations and realistic optical effects. Among these, image-based bokeh rendering has shown promising results, but diffusion for video bokeh remains unexplored. Existing image-based methods are plagued by temporal flickering and inconsistent blur transitions, while current video editing methods lack explicit control over the focus plane and bokeh intensity. These issues limit their applicability for controllable video bokeh. In this work, we propose a one-step diffusion framework for generating temporally coherent, depth-aware video bokeh rendering. The framework employs a multi-plane image (MPI) representation adapted to the focal plane to condition the video diffusion model, thereby enabling it to exploit strong 3D priors from pretrained backbones. To further enhance temporal stability, depth robustness, and detail preservation, we introduce a progressive training strategy. Experiments on synthetic and real-world benchmarks demonstrate superior temporal coherence, spatial accuracy, and controllability, outperforming prior baselines. This work represents the first dedicated diffusion framework for video bokeh generation, establishing a new baseline for temporally coherent and controllable depth-of-field effects. Project page is available at this website https://vivocameraresearch.github.io/any2bokeh/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 077e7157-caa3-42ab-bd88-4f657a7f6de8Cited by top-tier papers3
- Elastic3D: Controllable Stereo Video Conversion with Guided Latent DecodingNando Metzger, Prune Truong, Goutam Bhat, Konrad Schindler et al.CVPR 2026 · 3 citations
- Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic QualityZekai Luo, Zongze Du, Zhouhang Zhu, Hao Zhong et al.CVPR 2026 · 1 citation
- Towards Photorealistic and Efficient Bokeh Rendering via Diffusion FrameworkLinxiao Shi, Siming Zheng, Zerong Wang, Hao Zhang et al.CVPR 2026
Builds on14
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
- DragDiffusion: Harnessing Diffusion Models for Interactive Point-Based Image EditingYujun Shi, Chuhui Xue, Jun Hao Liew, Jiachun Pan et al.CVPR 2024 · 117 citations
- BokehMe: When Neural Rendering Meets Classical RenderingJuewen Peng, Zhiguo Cao, Xianrui Luo, Hao Lu et al.CVPR 2022 · 45 citations
- Dr.Bokeh: DiffeRentiable Occlusion-Aware Bokeh RenderingYichen Sheng, Zixun Yu, Lu Ling, Zhiwen Cao et al.CVPR 2024 · 9 citations
Related papers
- BokehCrafter: Taming Video Diffusion Models for Controllable Bokeh RenderingQiwen Wang, Liao Shen, Jiaqi Li, Tianqi Liu et al.AAAI 2026
- Video Bokeh Rendering: Make Casual Videography CinematicYawen Luo, Min Shi, Liao Shen, Yachuan Huang et al.ACM MM 2024 · 1 citation
- BokehFlow: Depth-Free Controllable Bokeh Rendering via Flow MatchingYachuan Huang, Xianrui Luo, Qiwen Wang, Liao Shen et al.AAAI 2026 · 2 citations
- UniScene-MoTion: Unified Scene & Motion-aware Diffusion Transition FrameworkRui Jiang, Chongmian Wang, Xinghe Fu, Yehao Lu et al.AAAI 2026
- BokehDiff: Neural Lens Blur with One-Step DiffusionChengxuan Zhu, Qingnan Fan, Qi Zhang, Jinwei Chen et al.ICCV 2025 · 2 citations
