TransFlow: Transformer as Flow Learner
Yawen Lu, Qifan Wang, Siqi Ma, Tong Geng, Yingjie Victor Chen, Huaijin G. Chen, Dongfang Liu
摘要
Optical flow is an indispensable building block for various important computer vision tasks, including motion estimation, object tracking, and disparity measurement. In this work, we propose TransFlow, a pure transformer architecture for optical flow estimation. Compared to dominant CNN-based methods, TransFlow demonstrates three advantages. First, it provides more accurate correlation and trustworthy matching in flow estimation by utilizing spatial self-attention and crossattention mechanisms between adjacent frames to effectively capture global dependencies; Second, it recovers more compromised information (e.g., occlusion and motion blur) in flow estimation through long-range temporal association in dynamic scenes; Third, it enables a concise self-learning paradigm and effectively eliminate the complex and laborious multi-stage pre-training procedures. We achieve the state-of-the-art results on the Sintel, KITTI-15, as well as several downstream tasks, including video object detection, interpolation and stabilization. For its efficacy, we hope TransFlow could serve as a flexible baseline for optical flow estimation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Recurrent Partial Kernel Network for Efficient Optical Flow EstimationHenrique Morimitsu, Xiaobin Zhu, Xiangyang Ji, Xu-Cheng YinAAAI 2024 · 被引用 29 次
- FlowDiffuser: Advancing Optical Flow Estimation with Diffusion ModelsAo Luo, Xin Li, Fan Yang, Jiangyu Liu 等CVPR 2024 · 被引用 26 次
- HDRFlow: Real-Time HDR Video Reconstruction with Large MotionsGangwei Xu, Yujin Wang, Jinwei Gu, Tianfan Xue 等CVPR 2024 · 被引用 12 次
- ProMotion: Prototypes as Motion LearnersYawen Lu, Dongfang Liu, Qifan Wang, Cheng Han 等CVPR 2024 · 被引用 8 次
- FlowSeek: Optical Flow Made Easier with Depth Foundation Models and Motion BasesMatteo Poggi, Fabio TosiICCV 2025 · 被引用 5 次
它引用的顶会 Paper27
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu 等ICCV 2021 · 被引用 2,462 次
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu 等ICCV 2021 · 被引用 2,397 次
- Uformer: A General U-Shaped Transformer for Image RestorationZhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou 等CVPR 2022 · 被引用 1,970 次
相关 Paper
- A Study of Finetuning Video Transformers for Multi-view Geometry TasksHuimin Wu, Kwang-Ting Cheng, Stephen Lin, Zhirong WuAAAI 2026
- CRAFT: Cross-Attentional Flow Transformer for Robust Optical FlowXiuchao Sui, Shaohua Li, Xue Geng, Yan Wu 等CVPR 2022 · 被引用 114 次
- Optical Flow Estimation from a Single Motion-blurred ImageDawit Mureja Argaw, Junsik Kim, François Rameau, Jae-Won Cho 等AAAI 2021 · 被引用 20 次
- High-Resolution Optical Flow from 1D Attention and CorrelationHaofei Xu, Jiaolong Yang, Jianfei Cai, Juyong Zhang 等ICCV 2021 · 被引用 92 次
- Learning to Estimate Hidden Motions with Global Motion AggregationShihao Jiang, Dylan Campbell, Yao Lu, Hongdong Li 等ICCV 2021 · 被引用 402 次
