MatchDet: A Collaborative Framework for Image Matching and Object Detection
Jinxiang Lai, Wenlong Wu, Bin-Bin Gao, Jun Liu, Jiawei Zhan, Congchong Nie, Yi Zeng, Chengjie Wang
Abstract
Image matching and object detection are two fundamental and challenging tasks, while many related applications consider them two individual tasks (i.e. task-individual). In this paper, a collaborative framework called MatchDet (i.e. task-collaborative) is proposed for image matching and object detection to obtain mutual improvements. To achieve the collaborative learning of the two tasks, we propose three novel modules, including a Weighted Spatial Attention Module (WSAM) for Detector, and Weighted Attention Module (WAM) and Box Filter for Matcher. Specifically, the WSAM highlights the foreground regions of target image to benefit the subsequent detector, the WAM enhances the connection between the foreground regions of pair images to ensure high-quality matches, and Box Filter mitigates the impact of false matches. We evaluate the approaches on a new benchmark with two datasets called Warp-COCO and miniScan-Net. Experimental results show our approaches are effective and achieve competitive improvements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b30c8051-7ee5-4e07-b88a-c60f56d3f272Builds on19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
Related papers
- Collaborative Feature Matching with Progressive Correspondence LearningXin Liu, Yanbing Han, Rong Qin, Bing Wang et al.AAAI 2026
- Weakly Supervised Object Detection With Segmentation CollaborationXiaoyan Li, Meina Kan, Shiguang Shan, Xilin ChenICCV 2019 · 105 citations
- Informative and Consistent Correspondence Mining for Cross-Domain Weakly Supervised Object DetectionLuwei Hou, Yu Zhang, Kui Fu, Jia LiCVPR 2021
- Contextually-Guided State Space Fusion for Misaligned Multi-Spectral Object DetectionGuyue Jin, Tianming Zhao, Jiacan Yan, Tian TianACM MM 2025 · 1 citation
- SynCL: A Synergistic Training Strategy with Instance-Aware Contrastive Learning for End-to-End Multi-Camera 3D TrackingShubo Lin, Yutong Kou, Zirui Wu, Shaoru Wang et al.NeurIPS 2025 · 2 citations
