Fuse2Match: Training-Free Fusion of Flow, Diffusion, and Contrastive Models for Zero-Shot Semantic Matching
Jing Zuo, Jiaqi Wang, Yonggang Qi, Yi-Zhe Song
摘要
Recent work shows that features from Stable Diffusion (SD) and contrastively pretrained models like DINO can be directly used for zero-shot semantic correspondence via naive feature concatenation. In this paper, we explore the stronger potential of Stable Diffusion 3 (SD3), a rectified flow-based model with a multimodal transformer backbone (MM-DiT). We show that semantic signals in SD3 are scattered across multiple timesteps and transformer layers, and propose a multi-level fusion scheme to extract discriminative features. Moreover, we identify that naive fusion across models suffers from inconsistent distributions, thus leading to suboptimal performance. To address this, we propose a simple yet effective confidence-aware feature fusion strategy that re-weights each model's contribution based on prediction confidence scores derived from their matching uncertainties. Notably, this fusion approach is not only training-free but also enables per-pixel adaptive integration of heterogeneous features. The resulting representation, Fuse2Match, significantly outperforms strong baselines on SPair-71k, PF-Pascal, and PSC6K, validating the benefit of combining SD3, SD, and DINO through our proposed confidence-aware feature fusion. Code is available at https://github.com/panda7777777/fuse2match
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic CorrespondenceJunyi Zhang, Charles Herrmann, Junhwa Hur, Luisa Polania Cabrera 等NeurIPS 2023 · 被引用 371 次
- Emergent Correspondence from Image DiffusionLuming Tang, Menglin Jia, Qianqian Wang, Cheng Perng Phoo 等NeurIPS 2023 · 被引用 555 次
- Bridging Viewpoint Gaps: Geometric Reasoning Boosts Semantic CorrespondenceQiyang Qian, Hansheng Chen, Masayoshi Tomizuka, Kurt Keutzer 等CVPR 2025
- Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive ActivationsChaofan Gan, Yuanpeng Tu, Xi Chen, Tieyuan Chen 等NeurIPS 2025 · 被引用 22 次
- Zero-shot RGB-D Point Cloud Registration with Pre-trained Large Vision ModelHaobo Jiang, Jin Xie, Jian Yang, Liang Yu 等CVPR 2025
