ARFlow: Auto-regressive Optical Flow Estimation for Arbitrary-Length Videos via Progressive Next-Frame Forecasting
Jiuming Liu, Mengmeng Liu, Siting Zhu, Yunpeng Zhang, Jiangtao Li, Michael Ying Yang, Francesco Nex, Hao Cheng, Hesheng Wang
摘要
Optical flow estimation is a fundamental computer vision task that predicts per-pixel displacements from consecutive images. Recent works attempt to exploit temporal cues to improve the estimation performance. However, their temporal modeling is restricted to short video sequences due to the unaffordable computational burden, thereby suffering from restricted temporal receptive fields. Moreover, their group-wise paradigm in one forward pass undermines inter-group information exchange, leading to modest performance improvement. To address these problems, we propose a novel multi-frame optical flow network based on an auto-regressive paradigm, named ARFlow. Unlike previous multi-frame methods, our method can be scalable to arbitrary-length videos with marginal computational overhead. Specifically, we design an Auto-regressive Flow Initialization (AFI) module and an Auto-regressive Multi-stride Flow Refinement (AMFR) module to forecast the next-frame flow based on multi-stride history observations. Our ARFlow achieves state-of-the-art performance, ranking 1st on both KITTI-2015 and Spring official benchmarks and 2nd on the MPI-Sintel (Final) benchmark among all open-sourced methods. Furthermore, due to the auto-regressive nature, our method can generalize to arbitrary video length with a constant GPU memory usage of 2.1GB.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- RL-ScanIQA: Reinforcement-Learned Scanpaths for Blind 360deg Image Quality AssessmentYujia Wang, Yuyan Li, Jiuming Liu, Fang-Lue Zhang 等CVPR 2026 · 被引用 3 次
- GoR: A Unified and Extensible Generative Framework for Ordinal RegressionHongxu Ma, Han Zhou, Kai Tian, Xuefeng Zhang 等ICLR 2026
- StreamVLO: Streaming Visual-LiDAR Odometry with Cumulative Drift CompensationMengmeng Liu, Jiuming Liu, Michael Ying Yang, Chaokang Jiang 等CVPR 2026
- UniFace: A fied ine-grained Understanding and Generation ModelJunzhe Li, Sifan Zhou, Liya Guo, Xuerui Qiu 等ICLR 2026
它引用的顶会 Paper48
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video DiffusionXun Huang, Zhengqi Li, Guande He, Mingyuan Zhou 等NeurIPS 2025 · 被引用 628 次
- Learning to Estimate Hidden Motions with Global Motion AggregationShihao Jiang, Dylan Campbell, Yao Lu, Hongdong Li 等ICCV 2021 · 被引用 402 次
- GMFlow: Learning Optical Flow via Global MatchingHaofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi 等CVPR 2022 · 被引用 353 次
- Incremental Transformer Structure Enhanced Image Inpainting with Masking Positional EncodingQiaole Dong, Chenjie Cao, Yanwei FuCVPR 2022 · 被引用 194 次
相关 Paper
- TransFlow: Transformer as Flow LearnerYawen Lu, Qifan Wang, Siqi Ma, Tong Geng 等CVPR 2023
- VideoFlow: Exploiting Temporal Cues for Multi-frame Optical Flow EstimationXiaoyu Shi, Zhaoyang Huang, Weikang Bian, Dasong Li 等ICCV 2023 · 被引用 112 次
- MemFlow: Optical Flow Estimation and Prediction with MemoryQiaole Dong, Yanwei FuCVPR 2024
- MEMFOF: High-Resolution Training for Memory-Efficient Multi-Frame Optical Flow EstimationVladislav Bargatin, Egor Chistov, Alexander Yakovenko, Dmitriy S. VatolinICCV 2025 · 被引用 11 次
- StreamFlow: Streamlined Multi-Frame Optical Flow Estimation for Video SequencesShangkun Sun, Jiaming Liu, Huaxia Li, Guoqing Liu 等NeurIPS 2024 · 被引用 19 次
