TransLO: A Window-Based Masked Point Transformer Framework for Large-Scale LiDAR Odometry
Jiuming Liu, Guangming Wang, Chaokang Jiang, Zhe Liu, Hesheng Wang
Abstract
Recently, transformer architecture has gained great success in the computer vision community, such as image classification, object detection, etc. Nonetheless, its application for 3D vision remains to be explored, given that point cloud is inherently sparse, irregular, and unordered. Furthermore, existing point transformer frameworks usually feed raw point cloud of N × 3 dimension into transformers, which limits the point processing scale because of their quadratic computational costs to the input size N . In this paper, we rethink the structure of point transformer. Instead of directly applying transformer to points, our network (TransLO) can process tens of thousands of points simultaneously by projecting points onto a 2D surface and then feeding them into a local transformer with linear complexity. Specifically, it is mainly composed of two components: Window-based Masked transformer with Self Attention (WMSA) to capture long-range dependencies; Masked Cross-Frame Attention (MCFA) to associate two frames and predict pose estimation. To deal with the sparsity issue of point cloud, we propose a binary mask to remove invalid and dynamic points. To our knowledge, this is the first transformer-based LiDAR odometry network. The experiment results on the KITTI odometry dataset show that our average rotational and translational RMSE achieves 0.500 • /100m and 0.993 % respectively. The performance of our network surpasses all recent learning-based methods and even outperforms LOAM on most evaluation sequences. Codes will be released on https://github.com/IRMVLab/TransLO .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 76680acb-7fef-40dc-9935-a86ce7d76315Cited by top-tier papers18
- RegFormer: An Efficient Projection-Aware Transformer Network for Large-Scale Point Cloud RegistrationJiuming Liu, Guangming Wang, Zhe Liu, Chaokang Jiang et al.ICCV 2023 · 71 citations
- Turboreg: Turboclique for Robust and Efficient Point Cloud RegistrationShaocheng Yan, Pengcheng Shi, Zhenjun Zhao, Kaixin Wang et al.ICCV 2025 · 11 citations
- Spherical Frustum Sparse Convolution Network for LiDAR Point Cloud Semantic SegmentationYu Zheng, Guangming Wang, Jiuming Liu, Marc Pollefeys et al.NeurIPS 2024 · 11 citations
- NeuroGauss4D-PCI: 4D Neural Fields and Gaussian Deformation Fields for Point Cloud InterpolationChaokang Jiang, Dalong Du, Jiuming Liu, Siting Zhu et al.NeurIPS 2024 · 10 citations
- S3E: Self-Supervised State Estimation for Radar-Inertial SystemShengpeng Wang, Yulong Xie, Qing Liao, Wei WangICCV 2025 · 3 citations
Builds on14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped WindowsXiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang et al.CVPR 2022 · 1,207 citations
- NICE-SLAM: Neural Implicit Scalable Encoding for SLAMZihan Zhu, Songyou Peng, Viktor Larsson, Weiwei Xu et al.CVPR 2022 · 720 citations
- Shunted Self-Attention via Multi-Scale Token AggregationSucheng Ren, Daquan Zhou, Shengfeng He, Jiashi Feng et al.CVPR 2022 · 326 citations
Related papers
- PWCLO-Net: Deep LiDAR Odometry in 3D Point Clouds Using Hierarchical Embedding Mask OptimizationGuangming Wang, Xinrui Wu, Zhe Liu, Hesheng WangCVPR 2021
- Voxel Set Transformer: A Set-to-Set Approach to 3D Object Detection from Point CloudsChenhang He, Ruihuang Li, Shuai Li, Lei ZhangCVPR 2022 · 217 citations
- SEFormer: Structure Embedding Transformer for 3D Object DetectionXiaoyu Feng, Heming Du, Hehe Fan, Yueqi Duan et al.AAAI 2023 · 15 citations
- Voxel Transformer for 3D Object DetectionJiageng Mao, Yujing Xue, Minzhe Niu, Haoyue Bai et al.ICCV 2021 · 535 citations
- SCTN: Sparse Convolution-Transformer Network for Scene Flow EstimationBing Li, Cheng Zheng, Silvio Giancola, Bernard GhanemAAAI 2022 · 50 citations
