Lune

ICCV2025顶会

Trackverse: a Large-Scale Object-Centric Video Dataset for Image-Level Representation Learning

Yibing Wei, Samuel Church, Victor Suciu, Jinhong Lin, Cheng-En Wu

2025年份
2被引次数

摘要

Video data inherently captures rich, dynamic contexts that reveal objects in varying poses, interactions, and state transitions, offering rich potential for unsupervised object representation learning. However, most prior representation learning methods rely on static image datasets like ImageNet, which lack temporal cues and only provide high-level semantic supervision. Meanwhile, existing natural video datasets are not ideal for learning objectcentric representations due to limited object focus and class diversity. To explore unsupervised object representation learning grounded in object dynamics-beyond static appearance-we introduce TrackVerse, a large-scale video dataset of 31.9 million object tracks spanning over 1,000 categories, each capturing the motion, appearance, and evolving states of an object over time. We further propose a variance-aware contrastive learning framework that adapts to data augmentations, encouraging the model to learn state-sensitive features. Extensive experiments demonstrate that representations learned from TrackVerse with variance-aware contrastive learning significantly outperform those from static image datasets and non-objectcentric natural video across multiple downstream tasks including object/attributie recognition, action recognition and video instance segmentation, highlighting the rich semantic and state content in TrackVerse feature.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper32

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖