Pixel-Level Semantic Correspondence Through Layout-Aware Representation Learning and Multi-Scale Matching Integration
Yixuan Sun, Zhangyue Yin, Haibo Wang, Yan Wang, Xipeng Qiu, Weifeng Ge, Wenqiang Zhang
Abstract
Establishing precise semantic correspondence across object instances in different images is a fundamental and challenging task in computer vision. In this task, difficulty arises often due to three challenges: confusing regions with similar appearance, inconsistent object scale, and indistinguishable nearby pixels. Recognizing these challenges, our paper proposes a novel semantic matching pipeline named LPMFlow toward extracting fine-grained semantics and geometry layouts for building pixel-level semantic correspondences. LPMFlow consists of three modules, each addressing one of the aforementioned challenges. The layout-aware representation learning module uniformly encodes source and target tokens to distinguish pixels or regions with similar appearances but different geometry semantics. The progressive feature superresolution module outputs four sets of 4D correlation tensors to generate accurate semantic flow between objects in different scales. Finally, the matching flow integration and refinement module is exploited to fuse matching flow in different scales to give the final flow predictions. The whole pipeline can be trained end-to-end, with a balance of computational cost and correspondence details. Extensive experiments based on benchmarks such as SPair-71K, PF-PASCAL, and PF-WILLOW have proved that the proposed method can well tackle the three challenges and outperform the previous methods, es-pecially in more stringent settings. Code is available at https://github.com/YXSUNMADMAX/LPMFlow.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8fce993-7c89-4169-9d2e-d9c9c6dc58beCited by top-tier papers2
- MARCO: Navigating the Unseen Space of Semantic CorrespondenceClaudia Cuttano, Gabriele Trivigno, Carlo Masone, Stefan RothCVPR 2026 · 4 citations
- AffordMatcher: Affordance Learning in 3D Scenes from Visual SignifiersNghia Vu, Tuong Do, Khang Nguyen, Baoru Huang et al.CVPR 2026 · 2 citations
Builds on23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment Everything Everywhere All at OnceXueyan Zou, Jianwei Yang, Hao Zhang, Feng Li et al.NeurIPS 2023 · 889 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with TransformersZhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy S. Ding et al.ICCV 2021 · 380 citations
Related papers
- Correspondence Transformers with Asymmetric Feature Learning and Matching Flow Super-ResolutionYixuan Sun, Dongyang Zhao, Zhangyue Yin, Yiwen Huang et al.CVPR 2023
- Multi-scale Matching Networks for Semantic CorrespondenceDongyang Zhao, Ziyang Song, Zhenghao Ji, Gangming Zhao et al.ICCV 2021 · 56 citations
- Dynamic Context Correspondence Network for Semantic AlignmentShuaiyi Huang, Qiuyue Wang, Songyang Zhang, Shipeng Yan et al.ICCV 2019 · 97 citations
- Hyperpixel Flow: Semantic Correspondence With Multi-Layer Neural FeaturesJuhong Min, Jongmin Lee, Jean Ponce, Minsu ChoICCV 2019 · 120 citations
- GLU-Net: Global-Local Universal Network for Dense Flow and CorrespondencesPrune Truong, Martin Danelljan, Radu TimofteCVPR 2020
