CATs: Cost Aggregation Transformers for Visual Correspondence
Seokju Cho, Sunghwan Hong, Sangryul Jeon, Yunsung Lee, Kwanghoon Sohn, Seungryong Kim
Abstract
We propose a novel cost aggregation network, called Cost Aggregation Transformers (CATs), to find dense correspondences between semantically similar images with additional challenges posed by large intra-class appearance and geometric variations. Cost aggregation is a highly important process in matching tasks, which the matching accuracy depends on the quality of its output. Compared to handcrafted or CNN-based methods addressing the cost aggregation, in that either lacks robustness to severe deformations or inherit the limitation of CNNs that fail to discriminate incorrect matches due to limited receptive fields, CATs explore global consensus among initial correlation map with the help of some architectural designs that allow us to fully leverage self-attention mechanism. Specifically, we include appearance affinity modeling to aid the cost aggregation process in order to disambiguate the noisy initial correlation maps and propose multi-level aggregation to efficiently capture different semantics from hierarchical feature representations. We then combine with swapping self-attention technique and residual connections not only to enforce consistent matching, but also to ease the learning process, which we find that these result in an apparent performance boost. We conduct experiments to demonstrate the effectiveness of the proposed model over the latest methods and provide extensive ablation studies. Code and trained models are available at https://sunghwanhong.github.io/CATs/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 332dd2b7-1883-45e7-b2f6-dacb55d560feCited by top-tier papers53
- Emergent Correspondence from Image DiffusionLuming Tang, Menglin Jia, Qianqian Wang, Cheng Perng Phoo et al.NeurIPS 2023 · 555 citations
- A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic CorrespondenceJunyi Zhang, Charles Herrmann, Junhwa Hur, Luisa Polania Cabrera et al.NeurIPS 2023 · 371 citations
- GAN-Supervised Dense Visual AlignmentWilliam S. Peebles, Jun-Yan Zhu, Richard Zhang, Antonio Torralba et al.CVPR 2022 · 50 citations
- Neural Matching Fields: Implicit Representation of Matching Fields for Visual CorrespondenceSunghwan Hong, Jisu Nam, Seokju Cho, Susung Hong et al.NeurIPS 2022 · 36 citations
- MSI: Maximize Support-Set Information for Few-Shot SegmentationSeonghyeon Moon, Samuel S. Sohn, Honglu Zhou, Sejong Yoon et al.ICCV 2023 · 34 citations
Builds on17
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
Related papers
- Unifying Feature and Cost Aggregation with Transformers for Semantic and Visual CorrespondenceSunghwan Hong, Seokju Cho, Seungryong Kim, Stephen LinICLR 2024 · 16 citations
- WT-MVSNet: Window-based Transformers for Multi-view StereoJinli Liao, Yikang Ding, Yoli Shavit, Dihe Huang et al.NeurIPS 2022 · 50 citations
- TransforMatcher: Match-to-Match Attention for Semantic CorrespondenceSeungwook Kim, Juhong Min, Minsu ChoCVPR 2022 · 26 citations
- Improving Transformer-based Image Matching by Cascaded Capturing Spatially Informative KeypointsChenjie Cao, Yanwei FuICCV 2023 · 23 citations
- CRFT: Consistent-Recurrent Feature Flow Transformer for Cross-Modal Image RegistrationXuecong Liu, Mengzhu Ding, Zixuan Sun, Zhang Li et al.CVPR 2026 · 4 citations
