RoMA: Scaling up Mamba-based Foundation Models for Remote Sensing
Fengxiang Wang, Yulin Wang, Mingshuo Chen, Haotian Wang, Hongzhen Wang, Haiyan Zhao, Yangang Sun, Shuo Wang, Di Wang, Long Lan, Wenjing Yang, Jing Zhang
摘要
Recent advances in self-supervised learning for Vision Transformers (ViTs) have fueled breakthroughs in remote sensing (RS) foundation models. However, the quadratic complexity of self-attention poses a significant barrier to scalability, particularly for large models and high-resolution images. While the linear-complexity Mamba architecture offers a promising alternative, existing RS applications of Mamba remain limited to supervised tasks on small, domain-specific datasets. To address these challenges, we propose RoMA, a framework that enables scalable self-supervised pretraining of Mamba-based RS foundation models using large-scale, diverse, unlabeled data. RoMA enhances scalability for high-resolution images through a tailored auto-regressive learning strategy, incorporating two key innovations: 1) a rotation-aware pretraining mechanism combining adaptive cropping with angular embeddings to handle sparsely distributed objects with arbitrary orientations, and 2) multi-scale token prediction objectives that address the extreme variations in object scales inherent to RS imagery. Systematic empirical studies validate that Mamba adheres to RS data and parameter scaling laws, with performance scaling reliably as model and data size increase. Furthermore, experiments across scene classification, object detection, and semantic segmentation tasks demonstrate that RoMA-pretrained Mamba models consistently outperform ViT-based counterparts in both accuracy and computational efficiency. The source code and pretrained models will be released at https://github.com/MiliLab/RoMA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and AnalysisZhengpeng Feng, Clement Atzberger, Sadiq Jaffer, Jovana Knezevic 等CVPR 2026 · 被引用 61 次
- SARMAE: Masked Autoencoder for SAR Representation LearningDanxu Liu, Di Wang, Hebaixu Wang, Haoyang Chen 等CVPR 2026 · 被引用 10 次
- Harnessing Massive Satellite Imagery with Efficient Masked Image ModelingFengxiang Wang, Hongzhen Wang, Di Wang, Zonghao Guo 等ICCV 2025 · 被引用 4 次
- ORSATR-X: A Foundation Model based on Differential-and-Excitation Networks for Optical Remote Sensing Object RecognitionCanyu Mo, Yongxiang Liu, Jiehua Zhang, Zilong Yu 等CVPR 2026 · 被引用 1 次
- XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?Fengxiang Wang, Hongzhen Wang, Zonghao Guo, Di Wang 等CVPR 2025
它引用的顶会 Paper33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
相关 Paper
- Cross-Scale MAE: A Tale of Multiscale Exploitation in Remote SensingMaofeng Tang, Andrei Cozma, Konstantinos Georgiou, Hairong QiNeurIPS 2023 · 被引用 88 次
- Mamba YOLO: A Simple Baseline for Object Detection with State Space ModelZeyu Wang, Chen Li, Huiying Xu, Xinzhong Zhu 等AAAI 2025 · 被引用 136 次
- ViT-Linearizer: Distilling Quadratic Knowledge into Linear-Time Vision ModelsGuoyizhe Wei, Rama ChellappaICCV 2025
- Mamba-Reg: Vision Mamba Also Needs RegistersFeng Wang, Jiahao Wang, Sucheng Ren, Guoyizhe Wei 等CVPR 2025
- T-APT: Text-Guided Modality-Aware Prompt Tuning for Arbitrary Multimodal Remote Sensing Data Joint ClassificationQinghao Gao, Jiahui Qu, Wenqian DongAAAI 2026
