Groupwise Query Specialization and Quality-Aware Multi-Assignment for Transformer-Based Visual Relationship Detection
Jongha Kim, Jihwan Park, Jinyoung Park, Jinyoung Kim, Sehyung Kim, Hyunwoo J. Kim
摘要
Visual Relationship Detection (VRD) has seen significant advancements with Transformer-based architectures recently. However, we identify two key limitations in a conventional label assignment for training Transformer-based VRD models, which is a process of mapping a ground-truth (GT) to a prediction. Under the conventional assignment, an 'unspecialized' query is trained since a query is expected to detect every relation, which makes it difficult for a query to specialize in specific relations. Furthermore, a query is also insufficiently trained since a GT is assigned only to a single prediction, therefore near-correct or even correct predictions are suppressed by being assigned 'no relation (∅)' as a GT. To address these issues, we propose Groupwise Query Specialization and Quality-Aware Multi-Assignment (SpeaQ). Groupwise Query Specialization trains a 'specialized' query by dividing queries and relations into disjoint groups and directing a query in a specific query group solely toward relations in the corresponding relation group. Quality-Aware Multi-Assignment further facilitates the training by assigning a GT to multiple predictions that are significantly close to a GT in terms of a subject, an object, and the relation in between. Experimental results and analyses show that SpeaQ effectively trains 'specialized' queries, which better utilize the capacity of a model, resulting in consistent performance gains with 'zero' additional inference cost across multiple VRD models and benchmarks. Code is available at https://github.com/mlvlab/SpeaQ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI DetectionChanhyeong Yang, Taehoon Song, Jihwan Park, Hyunwoo J. KimNeurIPS 2025 · 被引用 5 次
- Can We Build Scene Graphs, Not Classify Them? FlowSG: Progressive Image-Conditioned Scene Graph Generation with Flow MatchingXin Hu, Ke Qin, Wen Yin, Yuan-Fang Li 等CVPR 2026
- APT: Towards Universal Scene Graph Generation via Plug-in Adaptive Prompt TuningRuikun Luo, Changwei Gu, Jing Yang, Yuan Gao 等ICLR 2026
- Hydra-SGG: Hybrid Relation Assignment for One-stage Scene Graph GenerationMinghan Chen, Guikun Chen, Wenguan Wang, Yi YangICLR 2025
- RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction DetectionJihwan Park, Chanhyeong Yang, Jinyoung Park, Taehoon Song 等CVPR 2026
它引用的顶会 Paper27
- Group DETR: Fast DETR Training with Group-Wise One-to-Many AssignmentQiang Chen, Xiaokang Chen, Jian Wang, Shan Zhang 等ICCV 2023 · 被引用 231 次
- Mining the Benefits of Two-stage and One-stage HOI DetectionAixi Zhang, Yue Liao, Si Liu, Miao Lu 等NeurIPS 2021 · 被引用 218 次
- HOI Analysis: Integrating and Decomposing Human-Object InteractionYong-Lu Li, Xinpeng Liu, Xiaoqian Wu, Yizhuo Li 等NeurIPS 2020 · 被引用 152 次
- GEN-VLKT: Simplify Association and Enhance Interaction Understanding for HOI DetectionYue Liao, Aixi Zhang, Miao Lu, Yongliang Wang 等CVPR 2022 · 被引用 136 次
- Efficient Two-Stage Detection of Human-Object Interactions with a Novel Unary-Pairwise TransformerFrederic Z. Zhang, Dylan Campbell, Stephen GouldCVPR 2022 · 被引用 118 次
相关 Paper
- VRDFormer: End-to-End Video Visual Relation Detection with TransformersSipeng Zheng, Shizhe Chen, Qin JinCVPR 2022 · 被引用 16 次
- 3DVG-Transformer: Relation Modeling for Visual Grounding on Point CloudsLichen Zhao, Daigang Cai, Lu Sheng, Dong XuICCV 2021 · 被引用 234 次
- PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object DetectionZhengjian Kang, Jun Zhuang, Kangtong Mo, Qi Chen 等CVPR 2026 · 被引用 6 次
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou 等ICCV 2021 · 被引用 468 次
- DSGG: Dense Relation Transformer for an End-to-End Scene Graph GenerationZeeshan Hayder, Xuming HeCVPR 2024
