Unified Visual Relationship Detection with Vision and Language Models
Long Zhao, Liangzhe Yuan, Boqing Gong, Yin Cui, Florian Schroff, Ming-Hsuan Yang, Hartwig Adam, Ting Liu
Abstract
This work focuses on training a single visual relationship detector predicting over the union of label spaces from multiple datasets. Merging labels spanning different datasets could be challenging due to inconsistent taxonomies. The issue is exacerbated in visual relationship detection when second-order visual semantics are introduced between pairs of objects. To address this challenge, we propose UniVRD, a novel bottom-up method for Unified Visual Relationship Detection by leveraging vision and language models (VLMs). VLMs provide well-aligned image and text embeddings, where similar relationships are optimized to be close to each other for semantic unification. Our bottom-up design enables the model to enjoy the benefit of training with both object detection and visual relationship datasets. Empirical results on both human-object interaction detection and scene-graph generation demonstrate the competitive performance of our model. UniVRD achieves 38.07 mAP on HICO-DET, outperforming the current best bottom-up HOI detector by 14.26 mAP. More importantly, we show that our unified detector performs as well as dataset-specific models in mAP, and achieves further improvements when we scale up the model. Our code will be made publicly available on GitHub 1 . woman VG bed COCO sit on HICO lamp Obj365 look at V-COCO person COCO book Obj365 read HICO look at V-COCO streetlight Obj365 behind VG person COCO drive HICO bus COCO man VG wine glass Obj365 head VG part of VG person COCO hold V-COCO
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- RLIPv2: Fast Scaling of Relational Language-Image Pre-trainingHangjie Yuan, Shiwei Zhang, Xiang Wang, Samuel Albanie et al.ICCV 2023 · 69 citations
- Open-World Human-Object Interaction Detection via Multi-Modal PromptsJie Yang, Bingliang Li, Ailing Zeng, Lei Zhang et al.CVPR 2024 · 18 citations
- Toward Open-Set Human Object Interaction DetectionMingrui Wu, Yuqi Liu, Jiayi Ji, Xiaoshuai Sun et al.AAAI 2024 · 12 citations
- Learning Human-Object Interaction as GroupsJiajun Hong, Jianan Wei, Wenguan WangNeurIPS 2025 · 6 citations
- Towards Flexible Visual Relationship SegmentationFangrui Zhu, Jianwei Yang, Huaizu JiangNeurIPS 2024 · 4 citations
Builds on44
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation ModelsYichao Cao, Qingfei Tang, Xiu Su, Song Chen et al.NeurIPS 2023 · 64 citations
- Unifying 2D and 3D Vision-Language UnderstandingAyush Jain, Alexander Swerdlow, Yuzhou Wang, Sergio Arnaud et al.ICML 2025
- HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature AdaptationQinqian Lei, Bo Wang, Robby T. TanICCV 2025 · 4 citations
- GEN-VLKT: Simplify Association and Enhance Interaction Understanding for HOI DetectionYue Liao, Aixi Zhang, Miao Lu, Yongliang Wang et al.CVPR 2022 · 136 citations
- Discovering Syntactic Interaction Clues for Human-Object Interaction DetectionJinguo Luo, Weihong Ren, Weibo Jiang, Xi'ai Chen et al.CVPR 2024 · 10 citations
