Go-with-the-Track: Video Compositing and Motion Control with Point Tracking
Koichi Namekata, Yash Kant, Zhizheng Liu, Ryan D. Burgert, Yuancheng Xu, Kuan Heng Lin, Emmett Steven, Julien Philip, Li Ma, Andrea Vedaldi, Paul Debevec, Ning Yu
Abstract
Filmmaking demands precise motion control and reference image compositing -capabilities that existing methods treat separately. Point-trackconditioned image-to-video models restrict content insertion to the first frame, while reference-to-video models lack fine-grained spatial-temporal control over how reference content integrates across frames.
We present Go-with-the-Track, which unifies both capabilities by jointly conditioning on multiple reference images and reference-anchored point-tracks -extending conventional point-tracks to explicitly establish correspondences between generated frames and reference images, thus enabling precise compositing and motion control throughout the video.
To achieve this, we introduce spatially-aware point-track embeddings that encode the full sequence of point-track coordinates using a coordinate-wise MLP followed by temporal pooling. This representation captures the spatial characteristics of each point-track (serving as a unique identifier), while the embedding similarity correlates directly with spatial proximity, enhancing the model's ability to distinguish and associate point-tracks. We inject these point-track embeddings into a video diffusion transformer via a lightweight adapter, resolving the pixel-to-patch resolution mismatch while avoiding the substantial motion detail loss inherent in naive point-track subsampling.
We use a hybrid training strategy to train jointly on dynamic, static, and synthetic scene video datasets to boost motion controllability. Experiments demonstrate that Go-with-the-Track achieves superior motion and reference control in a single model and enables new capabilities: multireference conditioned video generation with point-track driven compositing, as well as camera control for both static and dynamic scenes. Project Page: https://eyeline-labs.github.io/Go-with-the-Track/ CCS Concepts: • Computing methodologies → Computer vision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9eb14752-c11c-4847-99fc-de98b7e59793Builds on50
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SAM 3: Segment Anything with ConceptsNicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath et al.ICLR 2026 · 1,103 citations
- Point Transformer V2: Grouped Vector Attention and Partition-based PoolingXiaoyang Wu, Yixing Lao, Li Jiang, Xihui Liu et al.NeurIPS 2022 · 924 citations
- TokenFlow: Consistent Diffusion Features for Consistent Video EditingMichal Geyer, Omer Bar-Tal, Shai Bagon, Tali DekelICLR 2024 · 439 citations
- π3: Permutation-Equivariant Visual Geometry LearningYifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang et al.ICLR 2026 · 318 citations
Related papers
- Generative Video Motion Editing with 3D Point TracksYao-Chih Lee, Zhoutong Zhang, Jiahui Huang, Jui-Hsien Wang et al.CVPR 2026 · 23 citations
- TrackGo: A Flexible and Efficient Method for Controllable Video GenerationHaitao Zhou, Chuang Wang, Rui Nie, Jinlin Liu et al.AAAI 2025 · 41 citations
- VD3D: Taming Large Video Diffusion Transformers for 3D Camera ControlSherwin Bahmani, Ivan Skorokhodov, Aliaksandr Siarohin, Willi Menapace et al.ICLR 2025
- Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video GenerationHyeonho Jeong, Chun-Hao P. Huang, Jong Chul Ye, Niloy J. Mitra et al.CVPR 2025
- Frame In-N-Out: Unbounded Controllable Image-to-Video GenerationBoyang Wang, Xuweiyi Chen, Matheus Gadelha, Zezhou ChengNeurIPS 2025 · 9 citations
