Exploring Long Tail Visual Relationship Recognition with Large Vocabulary
Sherif Abdelkarim, Aniket Agarwal, Panos Achlioptas, Jun Chen, Jiaji Huang, Boyang Li, Kenneth Church, Mohamed Elhoseiny
Abstract
Several approaches have been proposed in recent literature to alleviate the long-tail problem, mainly in object classification tasks. In this paper, we make the first largescale study concerning the task of Long-Tail Visual Relationship Recognition (LTVRR). LTVRR aims at improving the learning of structured visual relationships that come from the long-tail (e.g., "rabbit grazing on grass"). In this setup, the subject, relation, and object classes each follow a long-tail distribution. To begin our study and make a future benchmark for the community, we introduce two LTVRR-related benchmarks, dubbed VG8K-LT and GQA-LT, built upon the widely used Visual Genome and GQA datasets. We use these benchmarks to study the performance of several state-of-the-art long-tail models on the LTVRR setup. Lastly, we propose a visiolinguistic hubless (VilHub) loss and a Mixup augmentation technique adapted to LTVRR setup, dubbed as RelMix. Both VilHub and RelMix can be easily integrated on top of existing models and despite being simple, our results show that they can remarkably improve the performance, especially on tail classes. Benchmarks, code, and models have been made available at: https://github.com/Vision-CAIR/LTVRR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6431dc31-fea2-46ae-a0ad-dc967c17287dCited by top-tier papers8
- SGTR: End-to-end Scene Graph Generation with TransformerRongjie Li, Songyang Zhang, Xuming HeCVPR 2022 · 108 citations
- RelTransformer: A Transformer-Based Long-Tail Visual Relationship RecognitionJun Chen, Aniket Agarwal, Sherif Abdelkarim, Deyao Zhu et al.CVPR 2022 · 21 citations
- A Probabilistic Graphical Model Based on Neural-symbolic Reasoning for Visual Relationship DetectionDongran Yu, Bo Yang, Qianhao Wei, Anchen Li et al.CVPR 2022 · 18 citations
- Unified Visual Relationship Detection with Vision and Language ModelsLong Zhao, Liangzhe Yuan, Boqing Gong, Yin Cui et al.ICCV 2023 · 13 citations
- Single-Stage Visual Relationship Learning using Conditional QueriesAlakh Desai, Tz-Ying Wu, Subarna Tripathi, Nuno VasconcelosNeurIPS 2022 · 11 citations
Builds on3
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan et al.ICLR 2020 · 1,496 citations
- Equalization Loss for Long-Tailed Object RecognitionJingru Tan, Changbao Wang, Buyu Li, Quanquan Li et al.CVPR 2020
Related papers
- Leveraging Predicate and Triplet Learning for Scene Graph GenerationJiankai Li, Yunhong Wang, Xiefan Guo, Ruijie Yang et al.CVPR 2024
- LTGC: Long-Tail Recognition via Leveraging LLMs-Driven Generated ContentQihao Zhao, Yalun Dai, Hao Li, Wei Hu et al.CVPR 2024 · 22 citations
- Global and Local Mixture Consistency Cumulative Learning for Long-tailed Visual RecognitionsFei Du, Peng Yang, Qi Jia, Fengtao Nan et al.CVPR 2023
- Uniformly Distributed Category Prototype-Guided Vision-Language Framework for Long-Tail RecognitionXiaoxuan He, Siming Fu, Xinpeng Ding, Yuchen Cao et al.ACM MM 2023 · 6 citations
- Self-Supervised Relationship ProbingJiuxiang Gu, Jason Kuen, Shafiq R. Joty, Jianfei Cai et al.NeurIPS 2020 · 18 citations
