LitePT: Lighter Yet Stronger Point Transformer
Yuanwen Yue, Damien Robert, Jianyuan Wang, Sunghwan Hong, Jan Dirk Wegner, Christian Rupprecht, Konrad Schindler
摘要
Modern neural architectures for 3D point cloud processing contain both convolutional layers and attention blocks, but the best way to assemble them remains unclear. We analyse the role of different computational blocks in 3D point cloud networks and find an intuitive behaviour: convolution is adequate to extract low-level geometry at high-resolution in early layers, where attention is expensive without bringing any benefits; attention captures high-level semantics and context in low-resolution, deep layers more efficiently, where convolution inflates the parameter count. Guided by this design principle, we propose a new, improved 3D point cloud backbone that employs convolutions in early stages and switches to attention for deeper layers. To avoid the loss of spatial layout information when discarding redundant convolution layers, we introduce a novel, parameter-free 3D positional encoding, PointROPE. The resulting LitePT model has fewer parameters, runs faster, and uses less memory than the state-of-the-art Point Transformer V3, but nonetheless matches or outperforms it on a range of tasks and datasets. Code and models are available at: https://github.com/prs-eth/LitePT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper45
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui 等ICCV 2019 · 被引用 3,193 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu 等ICCV 2021 · 被引用 2,397 次
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 被引用 2,162 次
相关 Paper
- Positional Prompt Tuning for Efficient 3D Representation LearningShaochen Zhang, Zekun Qi, Runpei Dong, Xiuxiu Bai 等ACM MM 2025 · 被引用 2 次
- Improved MLP Point Cloud Processing with High-Dimensional Positional EncodingYanmei Zou, Hongshan Yu, Zhengeng Yang, Zechuan Li 等AAAI 2024 · 被引用 15 次
- On Geometry-Enhanced Parameter-Efficient Fine-Tuning for 3D Scene SegmentationLiyao Tang, Zhe Chen, Dacheng TaoNeurIPS 2025 · 被引用 5 次
- Cloud Transformers: A Universal Approach To Point Cloud Processing TasksKirill Mazur, Victor LempitskyICCV 2021 · 被引用 51 次
- Fast Point TransformerChunghyun Park, Yoonwoo Jeong, Minsu Cho, Jaesik ParkCVPR 2022
