Lune

ICML2025

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Hila Chefer, Uriel Singer, Amit Zohar, Yuval Kirstain, Adam Polyak, Yaniv Taigman, Lior Wolf, Shelly Sheynin

2025Year

Abstract

A ballet dancer twirls on the surface of a still lake at sunset" "A skateboarder performs jumps" "An acrobat executing a handstand on a narrow beam above water" "Fingers press into a shimmering slime ball" Figure 1. Text-to-video samples generated by VideoJAM. We present VideoJAM, a framework that explicitly instills a strong motion prior to any video generation model. Our framework significantly enhances motion coherence across a wide variety of motion types.