SOMA: Feature Gradient Enhanced Affine-Flow Matching for SAR-Optical Registration
Haodong Wang, Tao Zhuo, Xiuwei Zhang, Hanlin Yin, Wencong Wu, Yanning Zhang
摘要
Achieving pixel-level registration between SAR and optical images remains a challenging task due to their fundamentally different imaging mechanisms and visual characteristics. Although deep learning has achieved great success in many cross-modal tasks, its performance on SAR-Optical registration tasks is still unsatisfactory. Gradient-based information has traditionally played a crucial role in handcrafted descriptors by highlighting structural differences. However, such gradient cues have not been effectively leveraged in deep learning frameworks for SAR-Optical image matching. To address this gap, we propose SOMA, a dense registration framework that integrates structural gradient priors into deep features and refines alignment through a hybrid matching strategy. Specifically, we introduce the Feature Gradient Enhancer (FGE), which embeds multi-scale, multi-directional gradient filters into the feature space using attention and reconstruction mechanisms to boost feature distinctiveness. Furthermore, we propose the Global-Local Affine-Flow Matcher (GLAM), which combines affine transformation and flow-based refinement within a coarse-to-fine architecture to ensure both structural consistency and local accuracy. Experimental results demonstrate that SOMA significantly improves registration precision, increasing the CMR@1px by 12.29% on the SEN1-2 dataset and 18.50% on the GFGE_SO dataset. In addition, SOMA exhibits strong robustness and generalizes well across diverse scenes and resolutions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 被引用 769 次
- RFNet: Unsupervised Network for Mutually Reinforcing Multi-modal Image Registration and FusionHan Xu, Jiayi Ma, Jiteng Yuan, Zhuliang Le 等CVPR 2022 · 被引用 161 次
- Do Computer Vision Foundation Models Learn the Low-level Characteristics of the Human Visual System?Yancheng Cai, Fei Yin, Dounia Hammou, Rafal MantiukCVPR 2025
- OmniGlue: Generalizable Feature Matching with Foundation Model GuidanceHanwen Jiang, Arjun Karpur, Bingyi Cao, Qixing Huang 等CVPR 2024
相关 Paper
- CRFT: Consistent-Recurrent Feature Flow Transformer for Cross-Modal Image RegistrationXuecong Liu, Mengzhu Ding, Zixuan Sun, Zhang Li 等CVPR 2026 · 被引用 4 次
- Deep Algorithm Unrolling with Registration Embedding for PansharpeningTingting Wang, Yongxu Ye, Faming Fang, Guixu Zhang 等ACM MM 2023 · 被引用 8 次
- Multi-scale Matching Networks for Semantic CorrespondenceDongyang Zhao, Ziyang Song, Zhenghao Ji, Gangming Zhao 等ICCV 2021 · 被引用 56 次
- GOCor: Bringing Globally Optimized Correspondence Volumes into Your Neural NetworkPrune Truong, Martin Danelljan, Luc Van Gool, Radu TimofteNeurIPS 2020 · 被引用 89 次
- A Stepwise Matching Method for Multi-modal Image based on Cascaded NetworkJinming Mu, Shuiping Gou, Shasha Mao, Shankui ZhengACM MM 2021 · 被引用 5 次
