Rays as Pixels: Learning A Joint Distribution of Video and Camera Trajectories
Wonbong Jang, Shikun Liu, Soubhik Sanyal, Juan Perez, Kam Woh Ng, Sanskar Agrawal, Juan-Manuel Perez-Rua, Yiannis Douratsos, Tao Xiang
Abstract
Can we bridge the gap between perceiving camera trajectories and rendering novel views within a single generative framework? Recovering camera parameters from images and rendering scenes from novel viewpoints are considered the forward and inverse problems in the field of computer vision and graphics. Previous approaches treat these problems in isolation, often failing when image coverage is sparse or camera poses are ambiguous. In this work, we propose Rays as Pixels, a specialized Video Diffusion Model (VDM) that learns a joint distribution of videos and camera trajectories. We represent cameras as dense ray pixels (raxels) and simultaneously denoise them alongside video frames using a novel Decoupled Self-Cross Attention. This joint formulation enables us to: i) generate a video from multiple input images following a defined camera trajectory, ii) perform novel view synthesis from sparse views (without necessarily requiring camera poses), and iii) predict the camera trajectory from a raw video. We evaluate our model on pose estimation, camera-controlled video generation and validate its self-consistency. Please reference supplementary material for more qualitative results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9476522-b3d0-4f2c-aad8-03a4d7553b3fBuilds on28
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- Depth Anything 3: Recovering the Visual Space from Any ViewsHaotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen et al.ICLR 2026 · 720 citations
Related papers
- Cameras as Rays: Pose Estimation via Ray DiffusionJason Y. Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang et al.ICLR 2024 · 126 citations
- CameraSquad: Achieving Content Consistency in Parallel Multi-Trajectory Camera-Controlled Video GenerationZhufeng Xu, Xuan Gao, Bailin Deng, Yikang Ding et al.SIGGRAPH 2026
- Collaborative Video Diffusion: Consistent Multi-video Generation with Camera ControlZhengfei Kuang, Shengqu Cai, Hao He, Yinghao Xu et al.NeurIPS 2024 · 131 citations
- Look Beyond: Two-Stage Scene View Generation via Panorama and Video DiffusionXueyang Kang, Zhengkang Xiang, Zezheng Zhang, Kourosh KhoshelhamACM MM 2025
- Trajectory attention for fine-grained video motion controlZeqi Xiao, Wenqi Ouyang, Yifan Zhou, Shuai Yang et al.ICLR 2025
