SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations
Yunnan Wang, Kecheng Zheng, Jianyuan Wang, Minghao Chen, David Novotny, Christian Rupprecht, Yinghao Xu, Xing Zhu, Wenjun Zeng, Xin Jin, Yujun Shen
Abstract
The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both semantic and spatio-temporal information. While existing datasets have advanced either 3D understanding or video generation, a significant gap remains in providing a unified resource that supports both domains at scale. To bridge this chasm, we introduce SceneScribe-1M, a new large-scale, multi-modal video dataset. It comprises one million in-the-wild videos, each meticulously annotated with detailed textual descriptions, precise camera parameters, dense depth maps, and consistent 3D point tracks. We demonstrate the versatility and value of SceneScribe-1M by establishing benchmarks across a wide array of downstream tasks, including monocular depth estimation, scene reconstruction, and dynamic point tracking, as well as generative tasks such as text-to-video synthesis, with or without camera control. By open-sourcing SceneScribe-1M, we aim to provide a comprehensive benchmark and a catalyst for research, fostering the development of models that can both perceive the dynamic 3D world and generate controllable, realistic video content.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9176d2fd-b3cb-4028-993a-d79c57f3d505Cited by top-tier papers1
Ask how each one uses itBuilds on34
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category ReconstructionJeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone et al.ICCV 2021 · 686 citations
- Deep Patch Visual OdometryZachary Teed, Lahav Lipson, Jia DengNeurIPS 2023 · 323 citations
Related papers
- From an Image to a Scene: Learning to Imagine the World from a Million 360° VideosMatthew Wallingford, Anand Bhattad, Aditya Kusupati, Vivek Ramanujan et al.NeurIPS 2024 · 2 citations
- SpatialVID: A Large-Scale Video Dataset with Spatial AnnotationsJiahao Wang, Yufeng Yuan, Rujie Zheng, Youtian Lin et al.CVPR 2026 · 72 citations
- DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D VisionLu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao et al.CVPR 2024
- Free-Form Motion Control: Controlling the 6D Poses of Camera and Objects in Video GenerationXincheng Shuai, Henghui Ding, Zhenyuan Qin, Hao Luo et al.ICCV 2025 · 2 citations
- SynthVerse: A Large-Scale Diverse Synthetic Dataset for Point TrackingWeiguang Zhao, Haoran Xu, Xingyu Miao, Qin Zhao et al.SIGGRAPH 2026
