FUSER: Feed-Forward Multiview 3D Registration Transformer and SE(3)^N Diffusion Refinement
Haobo Jiang, Jin Xie, Jian Yang, Liang Yu, Jianmin Zheng
Abstract
Registration of multiview point clouds conventionally relies on extensive pairwise matching to build a pose graph for global synchronization, which is computationally expensive and inherently ill-posed without holistic geometric constraints. This paper proposes FUSER, the first feed-forward multiview registration transformer that jointly processes all scans in a unified, compact latent space to directly predict global poses without any pairwise estimation. To maintain tractability, FUSER encodes each scan into low-resolution superpoint features via a sparse 3D CNN that preserves absolute translation cues, and performs efficient intra-and inter-scan reasoning through a Geometric Alternating Attention module. Particularly, we transfer 2D attention priors from off-the-shelf foundation models to enhance 3D feature interaction and geometric consistency. Building upon FUSER, we further introduce FUSER-DF, an SE(3) N diffusion refinement framework to correct FUSER's estimates via denoising in the joint SE(3) N space. FUSER acts as a surrogate multiview registration model to construct the denoiser, and a prior-conditioned SE(3) N variational lower bound is derived for denoising supervision. Extensive experiments on 3DMatch, ScanNet and ArkitScenes demonstrate that our approach achieves the superior registration accuracy and outstanding computational efficiency. Code is available at https://github.com/Jiang-HB/FUSER.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on38
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
Related papers
- Dual Focus-Attention Transformer for Robust Point Cloud RegistrationKexue Fu, Mingzhi Yuan, Changwei Wang, Weiguang Pang et al.CVPR 2025
- Geometric Transformer for Fast and Robust Point Cloud RegistrationZheng Qin, Hao Yu, Changjian Wang, Yulan Guo et al.CVPR 2022 · 436 citations
- RGGT: A Generative-Prior-Guided Transformer for Unified Rigid and Non-Rigid Point Cloud RegistrationChengyu Zheng, Songlin Yang, Jin Huang, Honghua Chen et al.ICML 2026
- GGPT: Geometry-Grounded Point TransformerYutong Chen, Yiming Wang, Xucong Zhang, Sergey Prokudin et al.CVPR 2026 · 2 citations
- End2End Multi-View Feature Matching with Differentiable Pose OptimizationBarbara Roessle, Matthias NießnerICCV 2023 · 34 citations
