TransFlow: Transformer as Flow Learner
Yawen Lu, Qifan Wang, Siqi Ma, Tong Geng, Yingjie Victor Chen, Huaijin G. Chen, Dongfang Liu
Abstract
Optical flow is an indispensable building block for various important computer vision tasks, including motion estimation, object tracking, and disparity measurement. In this work, we propose TransFlow, a pure transformer architecture for optical flow estimation. Compared to dominant CNN-based methods, TransFlow demonstrates three advantages. First, it provides more accurate correlation and trustworthy matching in flow estimation by utilizing spatial self-attention and crossattention mechanisms between adjacent frames to effectively capture global dependencies; Second, it recovers more compromised information (e.g., occlusion and motion blur) in flow estimation through long-range temporal association in dynamic scenes; Third, it enables a concise self-learning paradigm and effectively eliminate the complex and laborious multi-stage pre-training procedures. We achieve the state-of-the-art results on the Sintel, KITTI-15, as well as several downstream tasks, including video object detection, interpolation and stabilization. For its efficacy, we hope TransFlow could serve as a flexible baseline for optical flow estimation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68e964fa-a65f-42f9-acb5-7131558711f3Cited by top-tier papers11
- Recurrent Partial Kernel Network for Efficient Optical Flow EstimationHenrique Morimitsu, Xiaobin Zhu, Xiangyang Ji, Xu-Cheng YinAAAI 2024 · 29 citations
- FlowDiffuser: Advancing Optical Flow Estimation with Diffusion ModelsAo Luo, Xin Li, Fan Yang, Jiangyu Liu et al.CVPR 2024 · 26 citations
- HDRFlow: Real-Time HDR Video Reconstruction with Large MotionsGangwei Xu, Yujin Wang, Jinwei Gu, Tianfan Xue et al.CVPR 2024 · 12 citations
- ProMotion: Prototypes as Motion LearnersYawen Lu, Dongfang Liu, Qifan Wang, Cheng Han et al.CVPR 2024 · 8 citations
- FlowSeek: Optical Flow Made Easier with Depth Foundation Models and Motion BasesMatteo Poggi, Fabio TosiICCV 2025 · 5 citations
Builds on27
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
- Uformer: A General U-Shaped Transformer for Image RestorationZhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou et al.CVPR 2022 · 1,970 citations
Related papers
- A Study of Finetuning Video Transformers for Multi-view Geometry TasksHuimin Wu, Kwang-Ting Cheng, Stephen Lin, Zhirong WuAAAI 2026
- CRAFT: Cross-Attentional Flow Transformer for Robust Optical FlowXiuchao Sui, Shaohua Li, Xue Geng, Yan Wu et al.CVPR 2022 · 114 citations
- Optical Flow Estimation from a Single Motion-blurred ImageDawit Mureja Argaw, Junsik Kim, François Rameau, Jae-Won Cho et al.AAAI 2021 · 20 citations
- High-Resolution Optical Flow from 1D Attention and CorrelationHaofei Xu, Jiaolong Yang, Jianfei Cai, Juyong Zhang et al.ICCV 2021 · 92 citations
- Learning to Estimate Hidden Motions with Global Motion AggregationShihao Jiang, Dylan Campbell, Yao Lu, Hongdong Li et al.ICCV 2021 · 402 citations
