SUDS: Scalable Urban Dynamic Scenes
Haithem Turki, Jason Y. Zhang, Francesco Ferroni, Deva Ramanan
Abstract
We extend neural radiance fields (NeRFs) to dynamic large-scale urban scenes. Prior work tends to reconstruct single video clips of short durations (up to 10 seconds). Two reasons are that such methods (a) tend to scale linearly with the number of moving objects and input videos because a separate model is built for each and (b) tend to require supervision via 3D bounding boxes and panoptic labels, obtained manually or via category-specific models. As a step towards truly open-world reconstructions of dynamic cities, we introduce two key innovations: (a) we factorize the scene into three separate hash table data structures to efficiently encode static, dynamic, and far-field radiance fields, and (b) we make use of unlabeled target signals consisting of RGB images, sparse LiDAR, off-the-shelf self-supervised 2D descriptors, and most importantly, 2D optical flow. Operationalizing such inputs via photometric, geometric, and feature-metric reconstruction losses enables SUDS to decompose dynamic scenes into the static background, individual objects, and their motions. When combined with our multi-branch table representation, such reconstructions can be scaled to tens of thousands of objects across 1.2 million frames from 1700 videos spanning geospatial footprints of hundreds of kilometers, (to our knowledge) the largest dynamic NeRF built to date. We present qualitative initial results on a variety of tasks enabled by our representations, including novel-view synthesis of dynamic urban scenes, unsupervised 3D instance segmentation, and unsupervised 3D cuboid detection. To compare to prior work, we also evaluate on KITTI and Virtual KITTI 2, surpassing state-ofthe-art methods that rely on ground truth 3D bounding box annotations while being 10x quicker to train.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85118440-d0c5-423b-8b0b-3d348573b7bcCited by top-tier papers64
- EmerNeRF: Emergent Spatial-Temporal Scene Decomposition via Self-SupervisionJiawei Yang, Boris Ivanovic, Or Litany, Xinshuo Weng et al.ICLR 2024 · 225 citations
- DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Autonomous Driving ScenesXiaoyu Zhou, Zhiwei Lin, Xiaojun Shan, Yongtao Wang et al.CVPR 2024 · 166 citations
- DOGS: Distributed-Oriented Gaussian Splatting for Large-Scale 3D Reconstruction Via Gaussian ConsensusYu Chen, Gim Hee LeeNeurIPS 2024 · 99 citations
- Cross-Ray Neural Radiance Fields for Novel-view Synthesis from Unconstrained Image CollectionsYifan Yang, Shuhai Zhang, Zixiong Huang, Yubing Zhang et al.ICCV 2023 · 61 citations
- NeuRAD: Neural Rendering for Autonomous DrivingAdam Tonderski, Carl Lindström, Georg Hess, William Ljungbergh et al.CVPR 2024 · 58 citations
Builds on30
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan et al.CVPR 2022 · 1,603 citations
- Plenoxels: Radiance Fields without Neural NetworksSara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen et al.CVPR 2022 · 1,237 citations
- Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields ReconstructionCheng Sun, Min Sun, Hwann-Tzong ChenCVPR 2022 · 859 citations
Related papers
- S-NeRF: Neural Radiance Fields for Street ViewsZiyang Xie, Junge Zhang, Wenye Li, Feihu Zhang et al.ICLR 2023 · 13 citations
- Unsupervised Multi-View Object Segmentation Using Radiance Field PropagationXinhang Liu, Jiaben Chen, Huai Yu, Yu-Wing Tai et al.NeurIPS 2022 · 34 citations
- Flux4D: Flow-based Unsupervised 4D ReconstructionJingkang Wang, Henry Che, Yun Chen, Ze Yang et al.NeurIPS 2025 · 10 citations
- D^2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular VideoTianhao Wu, Fangcheng Zhong, Andrea Tagliasacchi, Forrester Cole et al.NeurIPS 2022 · 184 citations
- SplatFlow: Self-Supervised Dynamic Gaussian Splatting in Neural Motion Flow Field for Autonomous DrivingSu Sun, Cheng Zhao, Zhuoyang Sun, Yingjie Victor Chen et al.CVPR 2025
