Multiscale Vision Transformers Meet Bipartite Matching for Efficient Single-Stage Action Localization
Ioanna Ntinou, Enrique Sanchez, Georgios Tzimiropoulos
摘要
In this paper, we observe that a straight bipartite matching loss can be applied to the output tokens of a vision transformer. This results in a backbone + MLP architecture that can do both tasks without the need of an extra encoder-decoder head and learnable queries. We show that a single MViTv2-S architecture trained with bipartite matching to perform both tasks surpasses the same MViTv2-S when trained with RoI align on pre-computed bounding boxes. With a careful design of token pooling and the proposed training pipeline, our Bipartite-Matching Vision Transformer model, BMViT, achieves +3 mAP on AVA2.2. w.r.t. the two-stage MViTv2-S counterpart. Code is available at https://github.com/IoannaNti/BMViT Action Localization is a challenging problem that combines detection and recognition tasks, which are often addressed separately. State-of-the-art methods rely on off-the-shelf bounding box detections pre-computed at high resolution, and propose transformer models that focus on the classification task alone. Such two-stage solutions are prohibitive for real-time deployment. On the other hand, single-stage methods target both tasks by devoting part of the network (generally the backbone) to sharing the majority of the work-load, compromising performance for speed. These methods build on adding a DETR head with learnable queries that after cross-and self-attention can be sent to corresponding MLPs for detecting a person's bounding box and action. However, DETR-like architectures are challenging to train and can incur in big complexity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li 等ICCV 2021 · 被引用 1,611 次
- MViTv2: Improved Multiscale Vision Transformers for Classification and DetectionYanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam 等CVPR 2022 · 被引用 699 次
相关 Paper
- PairDETR : Joint Detection and Association of Human Bodies and FacesAmmar Ali, Georgii Gaikov, Denis Rybalchenko, Alexander Chigorin 等CVPR 2024
- TxVAD: Improved Video Action Detection by TransformersZhenyu Wu, Zhou Ren, Yi Wu, Zhangyang Wang 等ACM MM 2022 · 被引用 5 次
- RegionViT: Regional-to-Local Attention for Vision TransformersChun-Fu Chen, Rameswar Panda, Quanfu FanICLR 2022 · 被引用 246 次
- WB-DETR: Transformer-Based Detector without BackboneFanfan Liu, Haoran Wei, Wenzhe Zhao, Guozhen Li 等ICCV 2021 · 被引用 45 次
- Efficient Two-Stage Detection of Human-Object Interactions with a Novel Unary-Pairwise TransformerFrederic Z. Zhang, Dylan Campbell, Stephen GouldCVPR 2022 · 被引用 118 次
