Seurat: From Moving Points to Depth
Seokju Cho, Jiahui Huang, Seungryong Kim, Joon-Young Lee
2025Year
8Top-tier citations
Abstract
Adobe Research 3D Trajectories Figure 1. Seurat predicts precise and smooth depth changes for dynamic objects over time by only looking at the 2D point trajectories, which encode depth cues in their motion patterns. The figure illustrates 2D point tracks lifted into 3D space with our depth predictions on videos from the DAVIS dataset [41].
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Emergent Temporal Correspondences from Video Diffusion TransformersJisu Nam, Soowon Son, Dahyun Chung, Jiyoung Kim et al.NeurIPS 2025 · 30 citations
- Generative Video Motion Editing with 3D Point TracksYao-Chih Lee, Zhoutong Zhang, Jiahui Huang, Jui-Hsien Wang et al.CVPR 2026 · 23 citations
- TrackingWorld: World-centric Monocular 3D Tracking of Almost All PixelsJiahao Lu, Weitao Xiong, Jiacheng Deng, Peng Li et al.NeurIPS 2025 · 7 citations
- AnthroTAP: Learning Point Tracking with Real-World MotionInès Hyeonsu Kim, Seokju Cho, Jahyeok Koo, Junghyun Park et al.CVPR 2026 · 5 citations
- DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry EstimationTuan Duc Ngo, Jiahui Huang, Seoung Wug Oh, Kevin Blackburn-Matzen et al.CVPR 2026 · 3 citations
Builds on22
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
Related papers
- SpatialTracker: Tracking Any 2D Pixels in 3D SpaceYuxi Xiao, Qianqian Wang, Shangzhan Zhang, Nan Xue et al.CVPR 2024 · 40 citations
- Swept volumes via spacetime numerical continuationSilvia Sellán, Noam Aigerman, Alec JacobsonSIGGRAPH 2021 · 35 citations
- DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular VideosWen-Hsuan Chu, Lei Ke, Katerina FragkiadakiNeurIPS 2024 · 75 citations
- AdaDepth: Exploiting Inherent Scene Information for Self-Supervised Depth Estimation in Dynamic ScenesXuanang Gao, Xiongbin Wu, Zhiwei Ning, Runze Yang et al.AAAI 2026
- SplineGS: Learning Smooth Trajectories in Gaussian Splatting for Dynamic Scene ReconstructionJihwan Yoon, Sangbeom Han, Jaeseok Oh, Minsik LeeICLR 2025
