RelTransformer: A Transformer-Based Long-Tail Visual Relationship Recognition
Jun Chen, Aniket Agarwal, Sherif Abdelkarim, Deyao Zhu, Mohamed Elhoseiny
摘要
The visual relationship recognition (VRR) task aims at understanding the pairwise visual relationships between interacting objects in an image. These relationships typically have a long-tail distribution due to their compositional nature. This problem gets more severe when the vocabulary becomes large, rendering this task very challenging. This paper shows that modeling an effective message-passing flow through an attention mechanism can be critical to tackling the compositionality and long-tail challenges in VRR. The method, called RelTransformer, represents each image as a fully-connected scene graph and restructures the whole scene into the relation-triplet and global-scene contexts. It directly passes the message from each element in the relation-triplet and global-scene contexts to the target relation via self-attention. We also design a learnable memory to augment the long-tail relation representation learning. Through extensive experiments, we find that our model generalizes well on many VRR benchmarks. Our model outperforms the best-performing models on two large-scale long-tail VRR benchmarks, VG8K-LT (+2.0% overall acc) and GQA-LT (+26.0% overall acc), both having a highly skewed distribution towards the tail. It also achieves strong results on the VG200 relation detection task. Our code is available at https://github.com/Vision-CAIR/ReITransformer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- With a Little Help from your own Past: Prototypical Memory Networks for Image CaptioningManuele Barraco, Sara Sarto, Marcella Cornia, Lorenzo Baraldi 等ICCV 2023 · 被引用 33 次
- DeiT-LT: Distillation Strikes Back for Vision Transformer Training on Long-Tailed DatasetsHarsh Rangwani, Pradipto Mondal, Mayank Mishra, Ashish Ramayee Asokan 等CVPR 2024 · 被引用 12 次
- Enhancing Masked Time-Series Modeling via Dropping PatchesTianyu Qiu, Yi Xie, Hao Niu, Yun Xiong 等AAAI 2025 · 被引用 3 次
- D2 Prune: Sparsifying Large Language Models via Dual Taylor Expansion and Attention Distribution AwarenessLang Xiong, Ning Liu, Ao Ren, Yuheng Bai 等AAAI 2026
- Leveraging Predicate and Triplet Learning for Scene Graph GenerationJiankai Li, Yunhong Wang, Xiefan Guo, Ruijie Yang 等CVPR 2024
它引用的顶会 Paper9
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan 等ICLR 2020 · 被引用 1,496 次
- On the Bottleneck of Graph Neural Networks and its Practical ImplicationsUri Alon, Eran YahavICLR 2021 · 被引用 90 次
- Exploring Long Tail Visual Relationship Recognition with Large VocabularySherif Abdelkarim, Aniket Agarwal, Panos Achlioptas, Jun Chen 等ICCV 2021 · 被引用 19 次
- Equalization Loss for Long-Tailed Object RecognitionJingru Tan, Changbao Wang, Buyu Li, Quanquan Li 等CVPR 2020
相关 Paper
- VRDFormer: End-to-End Video Visual Relation Detection with TransformersSipeng Zheng, Shizhe Chen, Qin JinCVPR 2022 · 被引用 16 次
- One-Shot Learning for Long-Tail Visual Relation DetectionWeitao Wang, Meng Wang, Sen Wang, Guodong Long 等AAAI 2020 · 被引用 20 次
- RetFormer: Multimodal Retrieval for Enhancing Image RecognitionTianrui Yu, Xiubo Liang, Hongzhi WangCVPR 2026
- Relation-Aware Graph Attention Network for Visual Question AnsweringLinjie Li, Zhe Gan, Yu Cheng, Jingjing LiuICCV 2019 · 被引用 391 次
- Unbiased Scene Graph Generation in VideosSayak Nag, Kyle Min, Subarna Tripathi, Amit K. Roy-ChowdhuryCVPR 2023
