BokehCrafter: Taming Video Diffusion Models for Controllable Bokeh Rendering
Qiwen Wang, Liao Shen, Jiaqi Li, Tianqi Liu, Huiqiang Sun, Zihao Huang, Yachuan Huang, Xianrui Luo, Zhiguo Cao
Abstract
Bokeh is used in photography to emphasize the selected subject by smoothly blurring the out-of-focus region with appealing highlights. While recent advances have achieved impressive results in rendering realistic blur, existing frameworks typically rely on disparity maps and bokeh-relevant inputs (e.g., focal distance and blur size), and face significant challenges in video bokeh rendering due to limited temporal consistency. In this paper, we propose BokehCrafter, the first video diffusion framework that generates temporally coherent and visually pleasing bokeh effects from all-in-focus video inputs under user-friendly input conditions. Specifically, we leverage a dual-stream attention mechanism, integrating a reference image branch and a rendering instruction branch. We propose a Bokeh Image Extraction (BIE) module and a CLIP-based text encoder to extract image and text features, respectively, whose outputs are fused via a Text-Image Fusion (TIF) module to enable fine-grained and controllable bokeh rendering. To support the novel capabilities of our model, we construct Video Bokeh Scenes (VBS), a large-scale dataset containing a wide variety of bokeh videos with corresponding rendering instructions, across various scenes and rendering settings. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art methods in both bokeh rendering quality and temporal consistency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Video Bokeh Rendering: Make Casual Videography CinematicYawen Luo, Min Shi, Liao Shen, Yachuan Huang et al.ACM MM 2024 · 1 citation
- Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion ModelYang Yang, Siming Zheng, Qirui Yang, Jinwei Chen et al.ICLR 2026 · 1 citation
- BokehFlow: Depth-Free Controllable Bokeh Rendering via Flow MatchingYachuan Huang, Xianrui Luo, Qiwen Wang, Liao Shen et al.AAAI 2026 · 2 citations
- BokehMe: When Neural Rendering Meets Classical RenderingJuewen Peng, Zhiguo Cao, Xianrui Luo, Hao Lu et al.CVPR 2022 · 45 citations
- Bokehlicious: Photorealistic Bokeh Rendering with Controllable AperturesTim Seizinger, Florin-Alexandru Vasluianu, Marcos V. Conde, Zongwei Wu et al.ICCV 2025 · 4 citations
