Matching Anything by Segmenting Anything
Siyuan Li, Lei Ke, Martin Danelljan, Luigi Piccinelli, Mattia Segù, Luc Van Gool, Fisher Yu
Abstract
The robust association of the same objects across video frames in complex scenes is crucial for many applications, especially multiple object tracking (MOT). Current methods predominantly rely on labeled domain-specific video datasets, which limits the cross-domain generalization of learned similarity embeddings. We propose MASA, a novel method for robust instance association learning, capable of matching any objects within videos across diverse domains without tracking labels. Leveraging the rich object segmentation from the Segment Anything Model (SAM), MASA learns instance-level correspondence through exhaustive data transformations. We treat the SAM outputs as dense object region proposals and learn to match those regions from a vast image collection. We further design a universal MASA adapter which can work in tandem with foundational segmentation or detection models and enable them to track any detected objects. Those combinations present strong zero-shot tracking ability in complex domains. Extensive tests on multiple challenging MOT and MOTS benchmarks indicate that the proposed method, using only unlabeled static images, achieves even better performance than stateof-the-art methods trained with fully annotated in-domain video sequences, in zero-shot association. Our code is available at github.com/siyuanliii/masa.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- DICEPTION: A Generalist Diffusion Model for Visual Perceptual TasksCanyu Zhao, Yanlong Sun, Mingyu Liu, Huanyi Zheng et al.NeurIPS 2025 · 45 citations
- UFM: A Simple Path towards Unified Dense Correspondence with FlowYuchen Zhang, Nikhil Varma Keetha, Chenwei Lyu, Bhuvan Jhamb et al.NeurIPS 2025 · 40 citations
- SAM2MOT: A Novel Paradigm of Multi-Object Tracking by SegmentationJunjie Jiang, Zelin Wang, Manqi Zhao, Yin Li et al.AAAI 2026 · 19 citations
- Language Decoupling with Fine-Grained Knowledge Guidance for Referring Multi-Object TrackingGuangyao Li, Siping Zhuang, Yajun Jian, Yan Yan et al.ICCV 2025 · 8 citations
- SegMASt3R: Geometry Grounded Segment MatchingRohit Jayanti, Swayam Agrawal, Vansh Garg, Siddharth Tourani et al.NeurIPS 2025 · 6 citations
Builds on32
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
Related papers
- SAM2-OV: A Novel Detection-Only Tuning Paradigm for Open-Vocabulary Multi-Object TrackingYangkai Chen, Qiangqiang Wu, Guangyao Li, Junlong Gao et al.AAAI 2026
- SAM2Object: Consolidating View Consistency via SAM2 for Zero-Shot 3D Instance SegmentationJihuai Zhao, Junbao Zhuo, Jiansheng Chen, Huimin MaCVPR 2025
- MV3DIS: Multi-View Mask Matching via 3D Guides for Zero-Shot 3D Instance SegmentationYibo Zhao, Yigong Zhang, Jin XieCVPR 2026 · 1 citation
- Segment Anything, Even OccludedWei-En Tai, Yu-Lin Shih, Cheng Sun, Yu-Chiang Frank Wang et al.CVPR 2025
- Towards Generalizable Scene Change DetectionJae-Woo Kim, Ue-Hwan KimCVPR 2025
