GMFlow: Learning Optical Flow via Global Matching
Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, Dacheng Tao
摘要
Learning-based optical flow estimation has been dominated with the pipeline of cost volume with convolutions for flow regression, which is inherently limited to local correlations and thus is hard to address the long-standing challenge of large displacements. To alleviate this, the state-of-the-art framework RAFT gradually improves its prediction quality by using a large number of iterative refinements, achieving remarkable performance but introducing linearly increasing inference time. To enable both high accuracy and efficiency, we completely revamp the dominant flow regression pipeline by reformulating optical flow as a global matching problem, which identifies the correspondences by directly comparing feature similarities. Specifically, we propose a GMFlow framework, which consists of three main components: a customized Transformer for feature enhancement, a correlation and softmax layer for global feature matching, and a self-attention layer for flow propagation. We further introduce a refinement step that reuses GMFlow at higher feature resolution for residual flow prediction. Our new framework outperforms 31-refinements RAFT on the challenging Sintel benchmark, while using only one refinement and running faster, suggesting a new paradigm for accurate and efficient optical flow estimation. Code is available at https://github.com/haofeixu/gmflow.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper146
- VideoComposer: Compositional Video Synthesis with Motion ControllabilityXiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen 等NeurIPS 2023 · 被引用 579 次
- Learning to Act from Actionless Videos through Dense CorrespondencesPo-Chen Ko, Jiayuan Mao, Yilun Du, Shao-Hua Sun 等ICLR 2024 · 被引用 181 次
- CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical FlowPhilippe Weinzaepfel, Thomas Lucas, Vincent Leroy, Yohann Cabon 等ICCV 2023 · 被引用 181 次
- The Surprising Effectiveness of Diffusion Models for Optical Flow and Monocular Depth EstimationSaurabh Saxena, Charles Herrmann, Junhwa Hur, Abhishek Kar 等NeurIPS 2023 · 被引用 160 次
- VideoFlow: Exploiting Temporal Cues for Multi-frame Optical Flow EstimationXiaoyu Shi, Zhaoyang Huang, Weikang Bian, Dasong Li 等ICCV 2023 · 被引用 112 次
它引用的顶会 Paper16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scale-Aware Trident Networks for Object DetectionYanghao Li, Yuntao Chen, Naiyan Wang, Zhaoxiang ZhangICCV 2019 · 被引用 1,031 次
- ConViT: Improving Vision Transformers with Soft Convolutional Inductive BiasesStéphane d'Ascoli, Hugo Touvron, Matthew L. Leavitt, Ari S. Morcos 等ICML 2021 · 被引用 1,021 次
- Perceiver IO: A General Architecture for Structured Inputs & OutputsAndrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch 等ICLR 2022 · 被引用 797 次
相关 Paper
- Global Matching with Overlapping Attention for Optical Flow EstimationShiyu Zhao, Long Zhao, Zhixing Zhang, Enyu Zhou 等CVPR 2022 · 被引用 85 次
- CRAFT: Cross-Attentional Flow Transformer for Robust Optical FlowXiuchao Sui, Shaohua Li, Xue Geng, Yan Wu 等CVPR 2022 · 被引用 114 次
- TransFlow: Transformer as Flow LearnerYawen Lu, Qifan Wang, Siqi Ma, Tong Geng 等CVPR 2023
- WAFT: Warping-Alone Field Transforms for Optical FlowYihan Wang, Jia DengICLR 2026 · 被引用 36 次
- DFlow: Learning to Synthesize Better Optical Flow Datasets via a Differentiable PipelineByung-Ki Kwon, Nam Hyeon-Woo, Ji-Yun Kim, Tae-Hyun OhICLR 2023
