Matching Every Pair to Track Every Point: PairFormer for All-Pairs Tracking and Video Trajectory Fields
Guangyang Wu, Youran Ding, Xinyu Che, Benyuan Sun, Yi Yang, Xiaohong Liu
Abstract
Tracking-any-point (TAP) answers query-conditioned correspondence but leaves the dense, all-pairs structure of a video implicit. We formulate All-Pairs Tracking (APT): given a video, predict dense displacement and visibility for every source-target frame pair, from which per-pixel trajectories can be read out. To this end, we propose Pair-Former, a feed-forward transformer that addresses APT in a single pass. A spatio-temporal patch encoder computes temporally conditioned features for all frames. CorrBank builds a learnable correlation bank for each frame pair and produces pairwise motion tokens. A broadcast motion mixer aggregates trajectory-wise context and broadcasts it back to refine the pairwise motion tokens. A trajectory head first predicts coarse dense displacement, visibility, and confidence, and then refines them iteratively to form a coherent all-pairs trajectory field. To support APT at scale, we develop PAIRender, a data platform that synthesizes photo-realistic dynamic scenes with dense annotations. From PAIRender we derive a training set (π-R10K) and a benchmark (APT-Bench) with an all-to-all evaluation protocol. Experiments show that PairFormer achieves strong performance on APT-Bench and competitive results on standard TAP benchmarks. Code and dataset will be released upon publication.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b0d4531-b783-4958-b60f-b543cdcaf79bBuilds on19
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- TAPIR: Tracking Any Point with per-frame Initialization and temporal RefinementCarl Doersch, Yi Yang, Mel Vecerík, Dilara Gokay et al.ICCV 2023 · 297 citations
- PointOdyssey: A Large-Scale Synthetic Dataset for Long-Term Point TrackingYang Zheng, Adam W. Harley, Bokui Shen, Gordon Wetzstein et al.ICCV 2023 · 255 citations
- Kubric: A scalable dataset generatorKlaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch et al.CVPR 2022 · 183 citations
- AccFlow: Backward Accumulation for Long-Range Optical FlowGuangyang Wu, Xiaohong Liu, Kunming Luo, Xi Liu et al.ICCV 2023 · 33 citations
Related papers
- AMT: All-Pairs Multi-Field Transforms for Efficient Frame InterpolationZhen Li, Zuo-Liang Zhu, Linghao Han, Qibin Hou et al.CVPR 2023
- CoWTracker: Tracking by Warping instead of CorrelationZihang Lai, Eldar Insafutdinov, Edgar Sucar, Andrea VedaldiCVPR 2026 · 12 citations
- TTAPFormer: Robust Arbitrary Point Tracking via Transient Asynchronous Fusion of Frames and EventsJiaxiong Liu, Zhen Tan, Jinpu Zhang, Yi Zhou et al.CVPR 2026
- Multi-View 3D Point TrackingFrano Rajic, Haofei Xu, Marko Mihajlovic, Siyuan Li et al.ICCV 2025 · 2 citations
- Learning to LEAP: Efficient Dense Point Tracking by Focusing Where It MattersChenzhi Zhao, Wufan Wang, Bo Zhang, Wendong WangAAAI 2026
