AKiRa: Augmentation Kit on Rays for Optical Video Generation
Xi Wang, Robin Courant, Marc Christie, Vicky Kalogeiton
2025Year
6Top-tier citations
Abstract
Figure 1. While current state-of-the-art video generation approaches offer limited control to users on camera motion, we propose a dedicated data augmentation framework -AKiRa -to train an optical video generation model that provides users with a panel of controls on camera motions (top row), camera focal length (second row), lens distortion (third row), or bokeh (camera aperture and focus in bottom row).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- EasyCreator: Empowering 4D Creation through Video InpaintingYue Ma, Kunyu Feng, Xinhua Zhang, Hongyu Liu et al.ICLR 2026 · 47 citations
- Unified Camera Positional Encoding for Controlled Video GenerationCheng Zhang, Boying Li, Meng Wei, Yan-Pei Cao et al.CVPR 2026 · 38 citations
- Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-DistillationSherwin Bahmani, Tianchang Shen, Jiawei Ren, Jiahui Huang et al.ICLR 2026 · 33 citations
- MotionCanvas: Cinematic Shot Design with Controllable Image-to-Video GenerationJinbo Xing, Long Mai, Cusuh Ham, Jiahui Huang et al.SIGGRAPH 2025 · 11 citations
- VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical PriorXindi Yang, Baolu Li, Yiming Zhang, Zhenfei Yin et al.ICCV 2025 · 8 citations
Builds on26
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Less is More: Data-Efficient Adaptation for Controllable Text-to-Video GenerationShihan Cheng, Nilesh Kulkarni, David Hyde, Dmitriy SmirnovCVPR 2026
- AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion TransformersSherwin Bahmani, Ivan Skorokhodov, Guocheng Qian, Aliaksandr Siarohin et al.CVPR 2025
- LeviTor: 3D Trajectory Oriented Image-to-Video SynthesisHanlin Wang, Hao Ouyang, Qiuyu Wang, Wen Wang et al.CVPR 2025
- FaceCam: Portrait Video Camera Control via Scale-Aware ConditioningWeijie Lyu, Ming-Hsuan Yang, Zhixin ShuCVPR 2026
- Learning to Generate Highly Dynamic Videos using Synthetic Motion DataWonjoon Jin, Jiyun Won, Janghyeok Han, Qi Dai et al.CVPR 2026
