TrajTok: Learning Trajectory Tokens Enhances Video Understanding
Chenhao Zheng, Jieyu Zhang, Jianing Zhang, Weikai Huang, Ashutosh Kumar, Quan Kong, Oncel Tuzel, Chun-Liang Li, Ranjay Krishna
Abstract
Tokenization in video models, typically through patchification, generates an excessive and redundant number of tokens. This severely limits video efficiency and scalability. While the recent trajectory-based tokenizers offer a promising solution by decoupling video duration from token count, they rely on complex, external segmentation and tracking pipelines that are slow and task-agnostic. We propose TrajTok, an end-to-end video tokenizer module that is fully integrated and co-trained with video models for a downstream objective, dynamically adapting its token granularity to semantic complexity, independent of video duration. TrajTok contains a unified segmenter that performs implicit clustering over pixels in both space and time to directly produce object trajectories in a single forward pass. By prioritizing downstream adaptability over pixel-perfect segmentation fidelity, TrajTok is lightweight, efficient, and yet empirically improves video understanding performance. With TrajTok, we implement a video CLIP model trained from scratch (TrajViT2). It achieves the best accuracy at scale across both classification and retrieval benchmarks, while maintaining efficiency comparable to the best token-merging methods. TrajTok also proves to be a versatile component beyond its role as a tokenizer. We show that it can be seamlessly integrated as either a probing head for pretrained visual features (TrajAdapter) or an alignment connector in vision–language models (TrajVLM) with especially strong performance in long-video reasoning.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- One Trajectory, One Token: Grounded Video Tokenization Via Panoptic Sub-Object TrajectoryChenhao Zheng, Jieyu Zhang, Mohammadreza Salehi, Ziqi Gao et al.ICCV 2025 · 7 citations
- VideoFlexTok: Flexible-Length Coarse-to-Fine Video TokenizationAndrei Atanov, Jesse Allardice, Roman Bachmann, Oğuzhan Fatih Kar et al.ICML 2026 · 3 citations
- Efficient Long Video Tokenization via Coordinate-based Patch ReconstructionHuiwon Jang, Sihyun Yu, Jinwoo Shin, Pieter Abbeel et al.CVPR 2025
- Unified Spatiotemporal Token Compression for Video-LLMs at Ultra-Low RetentionJunhao Du, Jialong Xue, Anqi Li, Jincheng Dai et al.CVPR 2026 · 7 citations
- VaporTok: RL-Driven Adaptive Video Tokenizer with Prior & Task AwarenessMinghao Yang, Zechen Bai, Jing Lin, Haoqian Wang et al.NeurIPS 2025 · 1 citation
