Lune

NeurIPS2023Top-tier venue

Language-based Action Concept Spaces Improve Video Self-Supervised Learning

Kanchana Ranasinghe, Michael S. Ryoo

2023Year
16Citations
6Top-tier citations

Abstract

Recent contrastive language image pre-training has led to learning highly transferable and robust image representations. However, adapting these models to video domain with minimal supervision remains an open problem. We explore a simple step in that direction, using language tied self-supervised learning to adapt an image CLIP model to the video domain. A backbone modified for temporal modeling is trained under self-distillation settings with train objectives operating in an action concept space. Feature vectors of various action concepts extracted from a language encoder using relevant textual prompts construct this space. A large language model aware of actions and their attributes generates the relevant textual prompts. We introduce two train objectives, concept distillation and concept alignment, that retain generality of original representations while enforcing relations between actions and their attributes. Our approach improves zero-shot and linear probing performance on three action recognition benchmarks.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext f930a8ce-70a3-48fc-97e7-9607ff0b2347

Cited by top-tier papers6

Ask how each one uses it

Builds on37

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines