DyST: Towards Dynamic Neural Scene Representations on Real-World Videos
Maximilian Seitzer, Sjoerd van Steenkiste, Thomas Kipf, Klaus Greff, Mehdi S. M. Sajjadi
摘要
Visual understanding of the world goes beyond the semantics and flat structure of individual images. In this work, we aim to capture both the 3D structure and dynamics of real-world scenes from monocular real-world videos. Our Dynamic Scene Transformer (DyST) model leverages recent work in neural scene representation to learn a latent decomposition of monocular real-world videos into scene content, per-view scene dynamics, and camera pose. This separation is achieved through a novel co-training scheme on monocular videos and our new synthetic dataset DySO. DyST learns tangible latent representations for dynamic scenes that enable view generation with separate control over the camera and the content of the scene.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion ModelsZiyi Wu, Yulia Rubanova, Rishabh Kabra, Drew A. Hudson 等NeurIPS 2024 · 被引用 32 次
- 4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any TimeZiqiao Ma, Xuweiyi Chen, Shoubin Yu, Sai Bi 等NeurIPS 2025 · 被引用 15 次
- Splat4D: Diffusion-Enhanced 4D Gaussian Splatting for Temporally and Spatially Consistent Content CreationMinghao Yin, Yukang Cao, Songyou Peng, Kai HanSIGGRAPH 2025 · 被引用 2 次
它引用的顶会 Paper17
- Dynamic View Synthesis from Dynamic Monocular VideoChen Gao, Ayush Saraf, Johannes Kopf, Jia-Bin HuangICCV 2021 · 被引用 522 次
- RegNeRF: Regularizing Neural Radiance Fields for View Synthesis from Sparse InputsMichael Niemeyer, Jonathan T. Barron, Ben Mildenhall, Mehdi S. M. Sajjadi 等CVPR 2022 · 被引用 513 次
- Light Field Networks: Neural Scene Representations with Single-Evaluation RenderingVincent Sitzmann, Semon Rezchikov, Bill Freeman, Josh Tenenbaum 等NeurIPS 2021 · 被引用 426 次
- Neural Radiance Flow for 4D View Synthesis and Video ProcessingYilun Du, Yinan Zhang, Hong-Xing Yu, Joshua B. Tenenbaum 等ICCV 2021 · 被引用 329 次
- NeRFPlayer: A Streamable Dynamic Scene Representation with Decomposed Neural Radiance FieldsLiangchen Song, Anpei Chen, Zhong Li, Zhang Chen 等IEEE VR 2023 · 被引用 246 次
相关 Paper
- Neural Scene Flow Fields for Space-Time View Synthesis of Dynamic ScenesZhengqi Li, Simon Niklaus, Noah Snavely, Oliver WangCVPR 2021
- Video Autoencoder: self-supervised disentanglement of static 3D structure and motionZihang Lai, Sifei Liu, Alexei A. Efros, Xiaolong WangICCV 2021 · 被引用 37 次
- Shape of Motion: 4D Reconstruction From a Single VideoQianqian Wang, Vickie Ye, Hang Gao, Weijia Zeng 等ICCV 2025 · 被引用 29 次
- DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular VideosWen-Hsuan Chu, Lei Ke, Katerina FragkiadakiNeurIPS 2024 · 被引用 75 次
- Unsupervised object-centric video generation and decomposition in 3DPaul Henderson, Christoph H. LampertNeurIPS 2020 · 被引用 41 次
