Lune

ACM MM2021Top-tier venue

Interventional Video Relation Detection

Yicong Li, Xun Yang, Xindi Shang, Tat-Seng Chua

2021Year
61Citations
23Top-tier citations

Abstract

Video Visual Relation Detection (VidVRD) aims to semantically describe the dynamic interactions across visual concepts localized in a video in the form of subject, predicate, object. It can help to mitigate the semantic gap between vision and language in video understanding, thus receiving increasing attention in multimedia communities. Existing efforts primarily leverage the multimodal/spatio-temporal feature fusion to augment the representation of object trajectories as well as their interactions and formulate the prediction of predicates as a multi-class classification task. Despite their effectiveness, existing models ignore the severe long-tailed bias in VidVRD datasets. As a result, the models' prediction will be easily biased towards the popular head predicates (e.g., next-to and in-front-of), thus leading to poor generalizability.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 5ab3d664-031f-4621-9628-3e77bfb7c514

Cited by top-tier papers23

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines