Lune

CVPR2026Top-tier venue

Matching Every Pair to Track Every Point: PairFormer for All-Pairs Tracking and Video Trajectory Fields

Guangyang Wu, Youran Ding, Xinyu Che, Benyuan Sun, Yi Yang, Xiaohong Liu

2026Year

Abstract

Tracking-any-point (TAP) answers query-conditioned correspondence but leaves the dense, all-pairs structure of a video implicit. We formulate All-Pairs Tracking (APT): given a video, predict dense displacement and visibility for every source-target frame pair, from which per-pixel trajectories can be read out. To this end, we propose Pair-Former, a feed-forward transformer that addresses APT in a single pass. A spatio-temporal patch encoder computes temporally conditioned features for all frames. CorrBank builds a learnable correlation bank for each frame pair and produces pairwise motion tokens. A broadcast motion mixer aggregates trajectory-wise context and broadcasts it back to refine the pairwise motion tokens. A trajectory head first predicts coarse dense displacement, visibility, and confidence, and then refines them iteratively to form a coherent all-pairs trajectory field. To support APT at scale, we develop PAIRender, a data platform that synthesizes photo-realistic dynamic scenes with dense annotations. From PAIRender we derive a training set (π-R10K) and a benchmark (APT-Bench) with an all-to-all evaluation protocol. Experiments show that PairFormer achieves strong performance on APT-Bench and competitive results on standard TAP benchmarks. Code and dataset will be released upon publication.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 3b0d4531-b783-4958-b60f-b543cdcaf79b

Builds on19

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines