Unifying Feature and Cost Aggregation with Transformers for Semantic and Visual Correspondence
Sunghwan Hong, Seokju Cho, Seungryong Kim, Stephen Lin
摘要
This paper introduces a Transformer-based integrative feature and cost aggregation network designed for dense matching tasks. In the context of dense matching, many works benefit from one of two forms of aggregation: feature aggregation, which pertains to the alignment of similar features, or cost aggregation, a procedure aimed at instilling coherence in the flow estimates across neighboring pixels. In this work, we first show that feature aggregation and cost aggregation exhibit distinct characteristics and reveal the potential for substantial benefits stemming from the judicious use of both aggregation processes. We then introduce a simple yet effective architecture that harnesses self- and cross-attention mechanisms to show that our approach unifies feature aggregation and cost aggregation and effectively harnesses the strengths of both techniques. Within the proposed attention layers, the features and cost volume both complement each other, and the attention layers are interleaved through a coarse-to-fine design to further promote accurate correspondence estimation. Finally at inference, our network produces multi-scale predictions, computes their confidence scores, and selects the most confident flow for final prediction. Our framework is evaluated on standard benchmarks for semantic matching, and also applied to geometric matching, where we show that our approach achieves significant improvements compared to existing methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Emergent Outlier View Rejection in Visual Geometry Grounded TransformersJisang Han, Sunghwan Hong, Jaewoo Jung, Wooseok Jang 等CVPR 2026 · 被引用 19 次
- PCA-Seg: Revisiting Cost Aggregation for Open-Vocabulary Semantic and Part SegmentationJianjian Yin, Tao Chen, Yi Chen, Gensheng Pei 等CVPR 2026 · 被引用 6 次
- S4M: Boosting Semi-Supervised Instance Segmentation with SAMHeeji Yoon, Heeseong Shin, Eunbeen Hong, Hyunwook Choi 等ICCV 2025 · 被引用 3 次
- Cross-View Completion Models are Zero-shot Correspondence EstimatorsHonggyu An, Jin Hyeon Kim, Seonghoon Park, Jaewoo Jung 等CVPR 2025
- Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part SegmentationJiho Choi, Seonho Lee, Minhyun Lee, Seungho Lee 等CVPR 2025
它引用的顶会 Paper25
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with TransformersZhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy S. Ding 等ICCV 2021 · 被引用 380 次
- COTR: Correspondence Transformer for Matching Across ImagesWei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi 等ICCV 2021 · 被引用 318 次
- CATs: Cost Aggregation Transformers for Visual CorrespondenceSeokju Cho, Sunghwan Hong, Sangryul Jeon, Yunsung Lee 等NeurIPS 2021 · 被引用 133 次
相关 Paper
- CoWTracker: Tracking by Warping instead of CorrelationZihang Lai, Eldar Insafutdinov, Edgar Sucar, Andrea VedaldiCVPR 2026 · 被引用 12 次
- Multi-Frame Self-Supervised Depth with TransformersVitor Guizilini, Rares Ambrus, Dian Chen, Sergey Zakharov 等CVPR 2022 · 被引用 95 次
- GOCor: Bringing Globally Optimized Correspondence Volumes into Your Neural NetworkPrune Truong, Martin Danelljan, Luc Van Gool, Radu TimofteNeurIPS 2020 · 被引用 89 次
- Multi-scale Matching Networks for Semantic CorrespondenceDongyang Zhao, Ziyang Song, Zhenghao Ji, Gangming Zhao 等ICCV 2021 · 被引用 56 次
- CRFT: Consistent-Recurrent Feature Flow Transformer for Cross-Modal Image RegistrationXuecong Liu, Mengzhu Ding, Zixuan Sun, Zhang Li 等CVPR 2026 · 被引用 4 次
