CRAFT: Cross-Attentional Flow Transformer for Robust Optical Flow
Xiuchao Sui, Shaohua Li, Xue Geng, Yan Wu, Xinxing Xu, Yong Liu, Rick Siow Mong Goh, Hongyuan Zhu
Abstract
Optical flow estimation aims to find the 2D motion field by identifying corresponding pixels between two images. Despite the tremendous progress of deep learning-based optical flow methods, it remains a challenge to accurately estimate large displacements with motion blur. This is mainly because the correlation volume, the basis of pixel matching, is computed as the dot product of the convolutional features of the two images. The locality of convolutional features makes the computed correlations susceptible to various noises. On large displacements with motion blur, noisy correlations could cause severe errors in the estimated flow. To overcome this challenge, we propose a new architecture “CRoss-Attentional Flow Trans-former” (CRAFT), aiming to revitalize the correlation volume computation. In CRAFT, a Semantic Smoothing Trans-former layer transforms the features of one frame, making them more global and semantically stable. In addition, the dot-product correlations are replaced with trans-former Cross-Frame Attention. This layer filters out feature noises through the Query and Key projections, and computes more accurate correlations. On Sintel (Final) and KITTI (foreground) benchmarks, CRAFT has achieved new state-of-the-art performance. Moreover, to test the robust-ness of different models on large motions, we designed an image shifting attack that shifts input images to generate large artificial motions. Under this attack, CRAFT per-forms much more robustly than two representative meth-ods, RAFT and GMA. The code of CRAFT is is available at https://github.com/askerlee/craft.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7e06574e-0970-43d2-a2b4-6b22e957a86aCited by top-tier papers39
- The Surprising Effectiveness of Diffusion Models for Optical Flow and Monocular Depth EstimationSaurabh Saxena, Charles Herrmann, Junhwa Hur, Abhishek Kar et al.NeurIPS 2023 · 160 citations
- VideoFlow: Exploiting Temporal Cues for Multi-frame Optical Flow EstimationXiaoyu Shi, Zhaoyang Huang, Weikang Bian, Dasong Li et al.ICCV 2023 · 112 citations
- UFM: A Simple Path towards Unified Dense Correspondence with FlowYuchen Zhang, Nikhil Varma Keetha, Chenwei Lyu, Bhuvan Jhamb et al.NeurIPS 2025 · 40 citations
- WAFT: Warping-Alone Field Transforms for Optical FlowYihan Wang, Jia DengICLR 2026 · 36 citations
- GAFlow: Incorporating Gaussian Attention into Optical FlowAo Luo, Fan Yang, Xin Li, Lang Nie et al.ICCV 2023 · 35 citations
Builds on14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- How Do Vision Transformers Work?Namuk Park, Songkuk KimICLR 2022 · 653 citations
- Self-supervised Co-Training for Video Representation LearningTengda Han, Weidi Xie, Andrew ZissermanNeurIPS 2020 · 405 citations
- Learning to Estimate Hidden Motions with Global Motion AggregationShihao Jiang, Dylan Campbell, Yao Lu, Hongdong Li et al.ICCV 2021 · 402 citations
Related papers
- GMFlow: Learning Optical Flow via Global MatchingHaofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi et al.CVPR 2022 · 353 citations
- TransFlow: Transformer as Flow LearnerYawen Lu, Qifan Wang, Siqi Ma, Tong Geng et al.CVPR 2023
- Global Matching with Overlapping Attention for Optical Flow EstimationShiyu Zhao, Long Zhao, Zhixing Zhang, Enyu Zhou et al.CVPR 2022 · 85 citations
- Learning Optical Flow with Kernel Patch AttentionAo Luo, Fan Yang, Xin Li, Shuaicheng LiuCVPR 2022 · 63 citations
- High-Resolution Optical Flow from 1D Attention and CorrelationHaofei Xu, Jiaolong Yang, Jianfei Cai, Juyong Zhang et al.ICCV 2021 · 92 citations
