SAMFlow: Eliminating Any Fragmentation in Optical Flow with Segment Anything Model
Shili Zhou, Ruian He, Weimin Tan, Bo Yan
摘要
Optical Flow Estimation aims to find the 2D dense motion field between two frames. Due to the limitation of model structures and training datasets, existing methods often rely too much on local clues and ignore the integrity of objects, resulting in fragmented motion estimation. Through theoretical analysis, we find the pre-trained large vision models are helpful in optical flow estimation, and we notice that the recently famous Segment Anything Model (SAM) demonstrates a strong ability to segment complete objects, which is suitable for solving the fragmentation problem. We thus propose a solution to embed the frozen SAM image encoder into FlowFormer to enhance object perception. To address the challenge of in-depth utilizing SAM in non-segmentation tasks like optical flow estimation, we propose an Optical Flow Task-Specific Adaption scheme, including a Context Fusion Module to fuse the SAM encoder with the optical flow context encoder, and a Context Adaption Module to adapt the SAM features for optical flow task with Learned Task-Specific Embedding. Our proposed SAMFlow model reaches 0.86/2.10 clean/final EPE and 3.55/12.32 EPE/F1-all on Sintel and KITTI-15 training set, surpassing Flowformer by 8.5%/9.9% and 13.2%/16.3%. Furthermore, our model achieves state-of-the-art performance on the Sintel and KITTI-15 benchmarks, ranking #1 among all two-frame methods on Sintel clean pass.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- WAFT: Warping-Alone Field Transforms for Optical FlowYihan Wang, Jia DengICLR 2026 · 被引用 36 次
- StreamFlow: Streamlined Multi-Frame Optical Flow Estimation for Video SequencesShangkun Sun, Jiaming Liu, Huaxia Li, Guoqing Liu 等NeurIPS 2024 · 被引用 19 次
- UnSAMFlow: Unsupervised Optical Flow Guided by Segment Anything ModelShuai Yuan, Lei Luo, Zhuo Hui, Can Pu 等CVPR 2024 · 被引用 7 次
- FacialFlowNet: Advancing Facial Optical Flow Estimation with a Diverse Dataset and a Decomposed ModelJianzhi Lu, Ruian He, Shili Zhou, Weimin Tan 等ACM MM 2024 · 被引用 4 次
- Context-Aware Iteration Policy Network for Efficient Optical Flow EstimationRi Cheng, Ruian He, Xuhao Jiang, Shili Zhou 等AAAI 2024 · 被引用 1 次
它引用的顶会 Paper16
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Perceiver IO: A General Architecture for Structured Inputs & OutputsAndrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch 等ICLR 2022 · 被引用 797 次
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang 等ICLR 2023 · 被引用 753 次
- GMFlow: Learning Optical Flow via Global MatchingHaofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi 等CVPR 2022 · 被引用 353 次
相关 Paper
- FlowFormer++: Masked Cost Volume Autoencoding for Pretraining Optical Flow EstimationXiaoyu Shi, Zhaoyang Huang, Dasong Li, Manyuan Zhang 等CVPR 2023
- A Study of Finetuning Video Transformers for Multi-view Geometry TasksHuimin Wu, Kwang-Ting Cheng, Stephen Lin, Zhirong WuAAAI 2026
- TransFlow: Transformer as Flow LearnerYawen Lu, Qifan Wang, Siqi Ma, Tong Geng 等CVPR 2023
- Efficient Track AnythingYunyang Xiong, Chong Zhou, Xiaoyu Xiang, Lemeng Wu 等ICCV 2025 · 被引用 5 次
- Learning Optical Flow with Kernel Patch AttentionAo Luo, Fan Yang, Xin Li, Shuaicheng LiuCVPR 2022 · 被引用 63 次
