Lune

ICCV2025Top-tier venue

Trackverse: a Large-Scale Object-Centric Video Dataset for Image-Level Representation Learning

Yibing Wei, Samuel Church, Victor Suciu, Jinhong Lin, Cheng-En Wu

2025Year
2Citations

Abstract

Video data inherently captures rich, dynamic contexts that reveal objects in varying poses, interactions, and state transitions, offering rich potential for unsupervised object representation learning. However, most prior representation learning methods rely on static image datasets like ImageNet, which lack temporal cues and only provide high-level semantic supervision. Meanwhile, existing natural video datasets are not ideal for learning objectcentric representations due to limited object focus and class diversity. To explore unsupervised object representation learning grounded in object dynamics-beyond static appearance-we introduce TrackVerse, a large-scale video dataset of 31.9 million object tracks spanning over 1,000 categories, each capturing the motion, appearance, and evolving states of an object over time. We further propose a variance-aware contrastive learning framework that adapts to data augmentations, encouraging the model to learn state-sensitive features. Extensive experiments demonstrate that representations learned from TrackVerse with variance-aware contrastive learning significantly outperform those from static image datasets and non-objectcentric natural video across multiple downstream tasks including object/attributie recognition, action recognition and video instance segmentation, highlighting the rich semantic and state content in TrackVerse feature.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Builds on32

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines