Groupwise Query Specialization and Quality-Aware Multi-Assignment for Transformer-Based Visual Relationship Detection
Jongha Kim, Jihwan Park, Jinyoung Park, Jinyoung Kim, Sehyung Kim, Hyunwoo J. Kim
Abstract
Visual Relationship Detection (VRD) has seen significant advancements with Transformer-based architectures recently. However, we identify two key limitations in a conventional label assignment for training Transformer-based VRD models, which is a process of mapping a ground-truth (GT) to a prediction. Under the conventional assignment, an 'unspecialized' query is trained since a query is expected to detect every relation, which makes it difficult for a query to specialize in specific relations. Furthermore, a query is also insufficiently trained since a GT is assigned only to a single prediction, therefore near-correct or even correct predictions are suppressed by being assigned 'no relation (∅)' as a GT. To address these issues, we propose Groupwise Query Specialization and Quality-Aware Multi-Assignment (SpeaQ). Groupwise Query Specialization trains a 'specialized' query by dividing queries and relations into disjoint groups and directing a query in a specific query group solely toward relations in the corresponding relation group. Quality-Aware Multi-Assignment further facilitates the training by assigning a GT to multiple predictions that are significantly close to a GT in terms of a subject, an object, and the relation in between. Experimental results and analyses show that SpeaQ effectively trains 'specialized' queries, which better utilize the capacity of a model, resulting in consistent performance gains with 'zero' additional inference cost across multiple VRD models and benchmarks. Code is available at https://github.com/mlvlab/SpeaQ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI DetectionChanhyeong Yang, Taehoon Song, Jihwan Park, Hyunwoo J. KimNeurIPS 2025 · 5 citations
- Can We Build Scene Graphs, Not Classify Them? FlowSG: Progressive Image-Conditioned Scene Graph Generation with Flow MatchingXin Hu, Ke Qin, Wen Yin, Yuan-Fang Li et al.CVPR 2026
- APT: Towards Universal Scene Graph Generation via Plug-in Adaptive Prompt TuningRuikun Luo, Changwei Gu, Jing Yang, Yuan Gao et al.ICLR 2026
- Hydra-SGG: Hybrid Relation Assignment for One-stage Scene Graph GenerationMinghan Chen, Guikun Chen, Wenguan Wang, Yi YangICLR 2025
- RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction DetectionJihwan Park, Chanhyeong Yang, Jinyoung Park, Taehoon Song et al.CVPR 2026
Builds on27
- Group DETR: Fast DETR Training with Group-Wise One-to-Many AssignmentQiang Chen, Xiaokang Chen, Jian Wang, Shan Zhang et al.ICCV 2023 · 231 citations
- Mining the Benefits of Two-stage and One-stage HOI DetectionAixi Zhang, Yue Liao, Si Liu, Miao Lu et al.NeurIPS 2021 · 218 citations
- HOI Analysis: Integrating and Decomposing Human-Object InteractionYong-Lu Li, Xinpeng Liu, Xiaoqian Wu, Yizhuo Li et al.NeurIPS 2020 · 152 citations
- GEN-VLKT: Simplify Association and Enhance Interaction Understanding for HOI DetectionYue Liao, Aixi Zhang, Miao Lu, Yongliang Wang et al.CVPR 2022 · 136 citations
- Efficient Two-Stage Detection of Human-Object Interactions with a Novel Unary-Pairwise TransformerFrederic Z. Zhang, Dylan Campbell, Stephen GouldCVPR 2022 · 118 citations
Related papers
- VRDFormer: End-to-End Video Visual Relation Detection with TransformersSipeng Zheng, Shizhe Chen, Qin JinCVPR 2022 · 16 citations
- 3DVG-Transformer: Relation Modeling for Visual Grounding on Point CloudsLichen Zhao, Daigang Cai, Lu Sheng, Dong XuICCV 2021 · 234 citations
- PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object DetectionZhengjian Kang, Jun Zhuang, Kangtong Mo, Qi Chen et al.CVPR 2026 · 6 citations
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou et al.ICCV 2021 · 468 citations
- DSGG: Dense Relation Transformer for an End-to-End Scene Graph GenerationZeeshan Hayder, Xuming HeCVPR 2024
