Easi3R: Estimating Disentangled Motion from DUSt3R Without Training
Xingyu Chen, Yue Chen, Yuliang Xiu, Andreas Geiger, Anpei Chen
摘要
Recent advances in DUSt3R have enabled robust estimation of dense point clouds and camera parameters of static scenes, leveraging Transformer network architectures and direct supervision on large-scale 3D datasets. In contrast, the limited scale and diversity of available 4D datasets present a major bottleneck for training a highly generalizable 4D model. This constraint has driven conventional 4D methods to fine-tune 3D models on scalable dynamic video data with additional geometric priors such as optical flow and depths. In this work, we take an opposite path and introduce Easi3R, a simple yet efficient training-free method for 4D reconstruction. Our approach applies attention adaptation during inference, eliminating the need for from-scratch pre-training or network fine-tuning. We find that the attention layers in DUSt3R inherently encode rich information about camera and object motion. By carefully disentangling these attention maps, we achieve accurate dynamic region segmentation, camera pose estimation, and 4D dense point map reconstruction. Extensive experiments on real-world dynamic videos demonstrate that our lightweight attention adaptation significantly outperforms previous state-of-the-art methods that are trained or finetuned on extensive dynamic datasets.
•
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper37
- TTT3R: 3D Reconstruction as Test-Time TrainingXingyu Chen, Yue Chen, Yuliang Xiu, Andreas Geiger 等ICLR 2026 · 被引用 139 次
- SpatialVID: A Large-Scale Video Dataset with Spatial AnnotationsJiahao Wang, Yufeng Yuan, Rujie Zheng, Youtian Lin 等CVPR 2026 · 被引用 72 次
- Efficiently Reconstructing Dynamic Scenes One D4RT at a TimeChuhan Zhang, Guillaume Le Moing, Skanda Koppula, Ignacio Rocco 等CVPR 2026 · 被引用 52 次
- Trace Anything: Representing Any Video in 4D via Trajectory FieldsXinhang Liu, Yuxi Xiao, Donny Y. Chen, Jiashi Feng 等ICLR 2026 · 被引用 38 次
- MoVieS: Motion-Aware 4D Dynamic View Synthesis in One SecondChenguo Lin, Yuchen Lin, Panwang Pan, Yifan Yu 等CVPR 2026 · 被引用 38 次
它引用的顶会 Paper31
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 被引用 1,248 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
- Prompt-to-Prompt Image Editing with Cross-Attention ControlAmir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman 等ICLR 2023 · 被引用 361 次
相关 Paper
- C4D: 4D Made from 3D Through Dual CorrespondencesShizun Wang, Zhenxiang Jiang, Xingyi Yang, Xinchao WangICCV 2025 · 被引用 4 次
- MonST3R: A Simple Approach for Estimating Geometry in the Presence of MotionJunyi Zhang, Charles Herrmann, Junhwa Hur, Varun Jampani 等ICLR 2025 · 被引用 3 次
- AnyCam: Learning to Recover Camera Poses and Intrinsics from Casual VideosFelix Wimbauer, Weirong Chen, Dominik Muhle, Christian Rupprecht 等CVPR 2025
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii 等CVPR 2024 · 被引用 302 次
- MoRe: Motion-aware Feed-forward 4D Reconstruction TransformerJuntong Fang, Zequn Chen, Weiqi Zhang, Donglin Di 等CVPR 2026 · 被引用 8 次
