Lune

NeurIPS2022Top-tier venue

Video-based Human-Object Interaction Detection from Tubelet Tokens

Danyang Tu, Wei Sun, Xiongkuo Min, Guangtao Zhai, Wei Shen

2022Year
24Citations
2Top-tier citations

Abstract

We present a novel vision Transformer, named TUTOR, which is able to learn tubelet tokens, served as highly-abstracted spatiotemporal representations, for video-based human-object interaction (V-HOI) detection. The tubelet tokens structurize videos by agglomerating and linking semantically-related patch tokens along spatial and temporal domains, which enjoy two benefits: 1) Compactness: each tubelet token is learned by a selective attention mechanism to reduce redundant spatial dependencies from others; 2) Expressiveness: each tubelet token is enabled to align with a semantic instance, i.e., an object or a human, across frames, thanks to agglomeration and linking. The effectiveness and efficiency of TUTOR are verified by extensive experiments. Results shows our method outperforms existing works by large margins, with a relative mAP gain of 16.14%16.14\% on VidHOI and a 2 points gain on CAD-120 as well as a 4×4 \times speedup.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 5e904f52-b0d3-4855-9aaf-8f546d23d678

Cited by top-tier papers2

Ask how each one uses it

Builds on23

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines