Object-Region Video Transformers
Roei Herzig, Elad Ben-Avraham, Karttikeya Mangalam, Amir Bar, Gal Chechik, Anna Rohrbach, Trevor Darrell, Amir Globerson
摘要
Recently, video transformers have shown great success in video understanding, exceeding CNN performance; yet existing video transformer models do not explicitly model objects, although objects can be essential for recognizing actions. In this work, we present Object-Region Video Transformers (ORViT), an object-centric approach that extends video transformer layers with a block that directly incorporates object representations. The key idea is to fuse object-centric representations starting from early layers and propagate them into the transformer-layers, thus affecting the spatio-temporal representations throughout the network. Our ORViT block consists of two object-level streams: appearance and dynamics. In the appearance stream, an “Object-Region Attention” module applies self-attention over the patches and object regions. In this way, visual object regions interact with uniform patch tokens and enrich them with contextualized object information. We further model object dynamics via a separate “Object-Dynamics Module”, which captures trajectory interactions, and show how to integrate the two streams. We evaluate our model on four tasks and five datasets: compositional and few-shot action recognition on SomethingElse, spatio-temporal action detection on AVA, and standard action recognition on Something-Something V2, Diving48 and Epic-Kitchen100. We show strong performance improvement across all tasks and datasets considered, demonstrating the value of a model that incorporates object representations into a transformer architecture. For code and pretrained models, visit the project page at https://roeiherz.github.io/ORViT/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- Masked Feature Prediction for Self-Supervised Visual Pre-TrainingChen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu 等CVPR 2022 · 被引用 524 次
- Omnivore: A Single Model for Many Visual ModalitiesRohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens van der Maaten 等CVPR 2022 · 被引用 185 次
- Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL ModelsSivan Doveh, Assaf Arbelle, Sivan Harary, Roei Herzig 等NeurIPS 2023 · 被引用 93 次
- Compositional Chain-of-Thought Prompting for Large Multimodal ModelsChancharik Mitra, Brandon Huang, Trevor Darrell, Roei HerzigCVPR 2024 · 被引用 62 次
- AIM: Adapting Image Models for Efficient Video Action RecognitionTaojiannan Yang, Yi Zhu, Yusheng Xie, Aston Zhang 等ICLR 2023 · 被引用 62 次
它引用的顶会 Paper20
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen 等NeurIPS 2020 · 被引用 3,042 次
相关 Paper
- BEVT: BERT Pretraining of Video TransformersRui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen 等CVPR 2022 · 被引用 200 次
- How can objects help action recognition?Xingyi Zhou, Anurag Arnab, Chen Sun, Cordelia SchmidCVPR 2023
- Deformable Video TransformerJue Wang, Lorenzo TorresaniCVPR 2022 · 被引用 40 次
- Keeping Your Eye on the Ball: Trajectory Attention in Video TransformersMandela Patrick, Dylan Campbell, Yuki M. Asano, Ishan Misra 等NeurIPS 2021 · 被引用 382 次
- Generative Video Transformer: Can Objects be the Words?Yi-Fu Wu, Jaesik Yoon, Sungjin AhnICML 2021 · 被引用 37 次
