ProMotion: Prototypes as Motion Learners
Yawen Lu, Dongfang Liu, Qifan Wang, Cheng Han, Yiming Cui, Zhiwen Cao, Xueling Zhang, Yingjie Victor Chen, Heng Fan
摘要
In this work, we introduce PRoMoTION, a unified proto-typical transformer-based framework engineered to model fundamental motion tasks. PRoMoTION offers a range of compelling attributes that set it apart from current task-specific paradigms. (1) We adopt a prototypical perspective, establishing a unified paradigm that harmonizes disparate motion learning approaches. This novel paradigm stream-lines the architectural design, enabling the simultaneous assimilation of diverse motion information. (2) We capitalize on a dual mechanism involving the feature denoiser and the prototypical learner to decipher the intricacies of motion. This approach effectively circumvents the pitfalls of ambiguity in pixel-wise feature matching, significantly bolstering the robustness of motion representation. (3)) We demon-strate a profound degree of transferability across distinct motion patterns. This inherent versatility reverberates robustly across a comprehensive spectrum of both 2D and 3D downstream tasks. Empirical results demonstrate that PRoMOTION outperforms various well-known specialized architectures, achieving 0.54 and 0.054 <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> error on the Sintel and KITTI depth datasets, 1.04 and 2.01 average endpoint error on the clean and final pass of Sintel flow benchmark, and 4.30 F1-all error on the KITTI flow bench-mark. For its efficacy, we hope our work can catalyze a paradigm shift in universal models in computer vision.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Prototypical Transformer As Unified Motion LearnersCheng Han, Yawen Lu, Guohao Sun, James Chenhao Liang 等ICML 2024 · 被引用 9 次
- Zero-Shot Compositional Video Learning with Coding Rate ReductionHeeseok Jung, Jun-Hyeon Bak, Yujin Jeong, Gyugeun Lee 等ICCV 2025 · 被引用 1 次
- PoseBH: Prototypical Multi-Dataset Training Beyond Human Pose EstimationUyoung Jeong, Jonathan Freer, Seungryul Baek, Hyung Jin Chang 等CVPR 2025
- Interpretable Image Classification via Non-parametric Part Prototype LearningZhijie Zhu, Lei Fan, Maurice Pagnucco, Yang SongCVPR 2025
它引用的顶会 Paper42
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped WindowsXiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang 等CVPR 2022 · 被引用 1,207 次
相关 Paper
- A Study of Finetuning Video Transformers for Multi-view Geometry TasksHuimin Wu, Kwang-Ting Cheng, Stephen Lin, Zhirong WuAAAI 2026
- TransFlow: Transformer as Flow LearnerYawen Lu, Qifan Wang, Siqi Ma, Tong Geng 等CVPR 2023
- MotionLab: Unified Human Motion Generation and Editing via the Motion-Condition-Motion ParadigmZiyan Guo, Zeyu Hu, De Wen Soh, Na ZhaoICCV 2025 · 被引用 10 次
- MultiMotion: Multi Subject Video Motion Transfer via Video Diffusion TransformerPenghui Liu, Jiangshan Wang, Yutong Shen, Shanhui Mo 等AAAI 2026 · 被引用 2 次
- MotionBERT: A Unified Perspective on Learning Human Motion RepresentationsWentao Zhu, Xiaoxuan Ma, Zhaoyang Liu, Libin Liu 等ICCV 2023 · 被引用 322 次
