TopicFM: Robust and Interpretable Topic-Assisted Feature Matching
Khang Truong Giang, Soohwan Song, Sungho Jo
Abstract
This study addresses an image-matching problem in challenging cases, such as large scene variations or textureless scenes. To gain robustness to such situations, most previous studies have attempted to encode the global contexts of a scene via graph neural networks or transformers. However, these contexts do not explicitly represent high-level contextual information, such as structural shapes or semantic instances; therefore, the encoded features are still not sufficiently discriminative in challenging scenes. We propose a novel image-matching method that applies a topic-modeling strategy to encode high-level contexts in images. The proposed method trains latent semantic instances called topics. It explicitly models an image as a multinomial distribution of topics, and then performs probabilistic feature matching. This approach improves the robustness of matching by focusing on the same semantic areas between the images. In addition, the inferred topics provide interpretability for matching the results, making our method explainable. Extensive experiments on outdoor and indoor datasets show that our method outperforms other state-of-the-art methods, particularly in challenging cases. Our code is available at github.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Efficient LoFTR: Semi-Dense Local Feature Matching with Sparse-Like SpeedYifan Wang, Xingyi He, Sida Peng, Dongli Tan et al.CVPR 2024 · 126 citations
- MESA: Matching Everything by Segmenting AnythingYesheng Zhang, Xu ZhaoCVPR 2024 · 19 citations
- EDM: Efficient Deep Feature MatchingXi Li, Tong Rao, Cihui PanICCV 2025 · 6 citations
- Semantic-aware Representation Learning for Homography EstimationYuhan Liu, Qianxin Huang, Siqi Hui, Jingwen Fu et al.ACM MM 2024 · 6 citations
- HOMO-Feature: Cross-Arbitrary-Modal Image Matching with Homomorphism of Organized Major OrientationChenzhong Gao, Wei Li, Desheng WengICCV 2025 · 3 citations
Builds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- DISK: Learning local features with policy gradientMichal J. Tyszkiewicz, Pascal Fua, Eduard TrullsNeurIPS 2020 · 652 citations
- Learning Two-View Correspondences and Geometry Using Order-Aware NetworkJiahui Zhang, Dawei Sun, Zixin Luo, Anbang Yao et al.ICCV 2019 · 362 citations
- COTR: Correspondence Transformer for Matching Across ImagesWei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi et al.ICCV 2021 · 318 citations
- Dual-Resolution Correspondence NetworksXinghui Li, Kai Han, Shuda Li, Victor PrisacariuNeurIPS 2020 · 207 citations
Related papers
- Scene-Aware Feature MatchingXiaoyong Lu, Yaping Yan, Tong Wei, Songlin DuICCV 2023 · 9 citations
- D2Former: Jointly Learning Hierarchical Detectors and Contextual Descriptors via Agent-Based TransformersJianfeng He, Yuan Gao, Tianzhu Zhang, Zhe Zhang et al.CVPR 2023
- TextFM: Robust Semi-dense Feature Matching with Language GuidanceZhihao Zheng, Jinglun Feng, Nirav Savaliya, Zheng-Hang Yeh et al.CVPR 2026 · 2 citations
- PMatch: Paired Masked Image Modeling for Dense Geometric MatchingShengjie Zhu, Xiaoming LiuCVPR 2023
- Co-Attention for Conditioned Image MatchingOlivia Wiles, Sébastien Ehrhardt, Andrew ZissermanCVPR 2021
