Direct Motion Models for Assessing Generated Videos
Kelsey R. Allen, Carl Doersch, Guangyao Zhou, Mohammed Suhail, Danny Driess, Ignacio Rocco, Yulia Rubanova, Thomas Kipf, Mehdi S. M. Sajjadi, Kevin Patrick Murphy, João Carreira, Sjoerd van Steenkiste
Abstract
A current limitation of video generative video models is that they generate plausible looking frames, but poor motion -an issue that is not well captured by FVD and other popular methods for evaluating generated videos. Here we go beyond FVD by developing a metric which better measures plausible object interactions and motion. Our novel approach is based on auto-encoding point tracks and yields motion features that can be used to not only compare distributions of videos (as few as one generated and one ground truth, or as many as two datasets), but also for evaluating motion of single videos. We show that using point tracks instead of pixel reconstruction or action recognition features results in a metric which is markedly more sensitive to temporal distortions in synthetic data, and can predict human evaluations of temporal consistency and realism in generated videos obtained from open-source models better than a wide range of alternatives. We also show that by using a point track representation, we can spatiotemporally localize generative video inconsistencies, providing extra interpretability of generated video errors relative to prior work. An overview of the results and link to the code can be found on the project page: trajan-paper.github.io. 1. At the distribution level, we show that among several other choices -including VideoMAE v2 (Wang et al., 2023b), I3D (Carreira and Zisserman, 2017), and motion his-Author contributions: KA conceptualized the idea of using track-based latent motion features to measure video quality, moving from distributional metrics to per-video or video-video metrics. KA, SVS co-led the project and ran all the experiments. GZ invented an initial prototype of the TRAJAN architecture, which was further developed by CD. MS supplied the WALT checkpoints. All authors advised on the project direction and contributed to writing the paper.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e1f5dfee-5d93-413b-aeee-2dd368b8d218Cited by top-tier papers4
- MoReGen: Multi-Agent Motion-Reasoning Engine for Code-based Text-to-Video SynthesisXiangyu Bai, He Liang, Bishoy Galoaa, Utsav Nandi et al.CVPR 2026 · 6 citations
- Dynamic Reflections: Probing Video Representations with Text AlignmentMaks Ovsjanikov, Viorica Patraucean, Leonidas J. Guibas, Tyler Zhu et al.ICLR 2026 · 5 citations
- Point Prompting: Counterfactual Tracking with Video Diffusion ModelsAyush Shrivastava, Sanyam Mehta, Daniel Geng, Andrew OwensICLR 2026 · 5 citations
- Learning Long-term Motion Embeddings for Efficient Kinematics GenerationNick Stracke, Kolja Bauer, Stefan Andreas Baumann, Miguel Ángel Bautista et al.CVPR 2026 · 2 citations
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- Perceiver: General Perception with Iterative AttentionAndrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals et al.ICML 2021 · 1,399 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video GenerationJay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei et al.ICCV 2023 · 1,113 citations
Related papers
- Beyond FVD: An Enhanced Evaluation Metrics for Video Generation Distribution QualityGe Ya Luo, Gian Mario Favero, Zhi Hao Luo, Alexia Jolicoeur-Martineau et al.ICLR 2025
- STREAM: Spatio-TempoRal Evaluation and Analysis Metric for Video Generative ModelsPum Jun Kim, Seojun Kim, Jaejun YooICLR 2024 · 11 citations
- EvalCrafter: Benchmarking and Evaluating Large Video Generation ModelsYaofang Liu, Xiaodong Cun, Xuebo Liu, Xintao Wang et al.CVPR 2024
- TrackMAE: Video Representation Learning via Track Mask and PredictRenaud Vandeghen, Fida Mohammad Thoker, Marc Van Droogenbroeck, Bernard GhanemCVPR 2026 · 3 citations
- On the Content Bias in Fréchet Video DistanceSongwei Ge, Aniruddha Mahapatra, Gaurav Parmar, Jun-Yan Zhu et al.CVPR 2024
