Trackverse: a Large-Scale Object-Centric Video Dataset for Image-Level Representation Learning
Yibing Wei, Samuel Church, Victor Suciu, Jinhong Lin, Cheng-En Wu
摘要
Video data inherently captures rich, dynamic contexts that reveal objects in varying poses, interactions, and state transitions, offering rich potential for unsupervised object representation learning. However, most prior representation learning methods rely on static image datasets like ImageNet, which lack temporal cues and only provide high-level semantic supervision. Meanwhile, existing natural video datasets are not ideal for learning objectcentric representations due to limited object focus and class diversity. To explore unsupervised object representation learning grounded in object dynamics-beyond static appearance-we introduce TrackVerse, a large-scale video dataset of 31.9 million object tracks spanning over 1,000 categories, each capturing the motion, appearance, and evolving states of an object over time. We further propose a variance-aware contrastive learning framework that adapts to data augmentations, encouraging the model to learn state-sensitive features. Extensive experiments demonstrate that representations learned from TrackVerse with variance-aware contrastive learning significantly outperform those from static image datasets and non-objectcentric natural video across multiple downstream tasks including object/attributie recognition, action recognition and video instance segmentation, highlighting the rich semantic and state content in TrackVerse feature.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- No More Shortcuts: Realizing the Potential of Temporal Self-SupervisionIshan Rajendrakumar Dave, Simon Jenni, Mubarak ShahAAAI 2024 · 被引用 14 次
- Rethinking Self-supervised Correspondence Learning: A Video Frame-level Similarity PerspectiveJiarui Xu, Xiaolong WangICCV 2021 · 被引用 112 次
- Temporally Consistent Object-Centric Learning by Contrasting SlotsAnna Manasyan, Maximilian Seitzer, Filip Radovic, Georg Martius 等CVPR 2025
- SeCo: Exploring Sequence Supervision for Unsupervised Representation LearningTing Yao, Yiheng Zhang, Zhaofan Qiu, Yingwei Pan 等AAAI 2021 · 被引用 118 次
- Contrastive Learning of Image Representations with Cross-Video Cycle-ConsistencyHaiping Wu, Xiaolong WangICCV 2021 · 被引用 35 次
