UFM: A Simple Path towards Unified Dense Correspondence with Flow
Yuchen Zhang, Nikhil Varma Keetha, Chenwei Lyu, Bhuvan Jhamb, Yutian Chen, Yuheng Qiu, Jay Karhade, Shreyas Jha, Yaoyu Hu, Deva Ramanan, Sebastian A. Scherer, Wenshan Wang
Abstract
Dense image correspondence is central to many applications, such as visual odometry, 3D reconstruction, object association, and re-identification. Historically, dense correspondence has been tackled separately for wide-baseline scenarios and optical flow estimation, despite the common goal of matching content between two images. In this paper, we develop a Unified Flow & Matching model (UFM), which is trained on unified data for pixels that are co-visible in both source and target images. UFM uses a simple, generic transformer architecture that directly regresses the (u, v) flow. It is easier to train and more accurate for large flows compared to the typical coarse-to-fine cost volumes in prior work. UFM is 28% more accurate than state-of-the-art flow methods (Unimatch), while also having 62% less error and 6.7x faster than dense wide-baseline matchers (RoMa). UFM is the first to demonstrate that unified training can outperform specialized approaches across both domains. This result enables fast, general-purpose correspondence and opens new directions for multi-modal, long-range, and real-time correspondence tasks.
39th Conference on Neural Information Processing Systems (NeurIPS 2025).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab0c4ffd-11d5-44c9-a1af-359c92b1985fCited by top-tier papers10
- Any4D: Unified Feed-Forward Metric 4D ReconstructionJay Karhade, Nikhil Varma Keetha, Yuchen Zhang, Tanisha Gupta et al.CVPR 2026 · 35 citations
- E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-trainingQitao Zhao, Hao Tan, Qianqian Wang, Sai Bi et al.CVPR 2026 · 24 citations
- In Pursuit of Pixel Supervision for Visual Pre-trainingLihe Yang, Shang-Wen Li, Yang Li, Xinjie Lei et al.CVPR 2026 · 13 citations
- PhysGM: Large Physical Gaussian Model for Feed-Forward 4D SynthesisChunji Lv, Zequn Chen, Donglin Di, Weinan Zhang et al.CVPR 2026 · 9 citations
- Flow3r: Factored Flow Prediction for Scalable Visual Geometry LearningZhongxiao Cong, Qitao Zhao, Minsik Jeon, Shubham TulsianiCVPR 2026 · 8 citations
Builds on30
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- LightGlue: Local Feature Matching at Light SpeedPhilipp Lindenberger, Paul-Edouard Sarlin, Marc PollefeysICCV 2023 · 936 citations
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans et al.NeurIPS 2021 · 826 citations
Related papers
- COTR: Correspondence Transformer for Matching Across ImagesWei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi et al.ICCV 2021 · 318 citations
- MV-RoMa: From Pairwise Matching into Multi-View Track ReconstructionJongMin Lee, Seungyeop Kang, Sungjoo YooCVPR 2026 · 3 citations
- PMatch: Paired Masked Image Modeling for Dense Geometric MatchingShengjie Zhu, Xiaoming LiuCVPR 2023
- UniCorrn: Unified Correspondence Transformer Across 2D and 3DPrajnan Goswami, Tianye Ding, Feng Liu, Huaizu JiangCVPR 2026 · 2 citations
- GLU-Net: Global-Local Universal Network for Dense Flow and CorrespondencesPrune Truong, Martin Danelljan, Radu TimofteCVPR 2020
