Consistency Learning via Decoding Path Augmentation for Transformers in Human Object Interaction Detection
Jihwan Park, Seungjun Lee, Hwan Heo, Hyeong Kyu Choi, Hyunwoo J. Kim
Abstract
Human-Object Interaction detection is a holistic visual recognition task that entails object detection as well as interaction classification. Previous works of HOI detection has been addressed by the various compositions of subset predictions, e.g., Image → HO → I, Image → HI → <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> . Recently, transformer based architecture for HOI has emerged, which directly predicts the HOI triplets in an end-to-end fashion (Image → HOI). Motivated by various inference paths for HOI detection, we propose cross-path consistency learning (CPC), which is a novel end-to-end learning strategy to improve HOI detection for transformers by leveraging augmented decoding paths. CPC learning enforces all the possible predictions from permuted inference sequences to be consistent. This simple scheme makes the model learn consistent representations, thereby improving generalization without increasing model capacity. Our experiments demonstrate the effectiveness of our method, and we achieved significant improvement on V-COCO and HICO-DET compared to the baseline models. Our code is available at https://github.com/mlvlab/CPChoi.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da10b9c8-5a2f-4076-8375-a37c1946fc7eCited by top-tier papers13
- RLIPv2: Fast Scaling of Relational Language-Image Pre-trainingHangjie Yuan, Shiwei Zhang, Xiang Wang, Samuel Albanie et al.ICCV 2023 · 69 citations
- TokenMixup: Efficient Attention-guided Token-level Data Augmentation for TransformersHyeong Kyu Choi, Joonmyung Choi, Hyunwoo J. KimNeurIPS 2022 · 49 citations
- Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion ModelsLiulei Li, Wenguan Wang, Yi YangNeurIPS 2024 · 29 citations
- Unified Visual Relationship Detection with Vision and Language ModelsLong Zhao, Liangzhe Yuan, Boqing Gong, Yin Cui et al.ICCV 2023 · 13 citations
- Open Set Video HOI detection from Action-centric Chain-of-Look PromptingNan Xi, Jingjing Meng, Junsong YuanICCV 2023 · 10 citations
Builds on24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
Related papers
- End-to-End Human Object Interaction Detection With HOI TransformerCheng Zou, Bohan Wang, Yue Hu, Junqi Liu et al.CVPR 2021
- What to look at and where: Semantic and Spatial Refined Transformer for detecting human-object interactionsA. S. M. Iftekhar, Hao Chen, Kaustav Kundu, Xinyu Li et al.CVPR 2022 · 50 citations
- HOTR: End-to-End Human-Object Interaction Detection With TransformersBumsoo Kim, Junhyun Lee, Jaewoo Kang, Eun-Sol Kim et al.CVPR 2021
- Efficient Two-Stage Detection of Human-Object Interactions with a Novel Unary-Pairwise TransformerFrederic Z. Zhang, Dylan Campbell, Stephen GouldCVPR 2022 · 118 citations
- MSTR: Multi-Scale Transformer for End-to-End Human-Object Interaction DetectionBumsoo Kim, Jonghwan Mun, Kyoung-Woon On, Minchul Shin et al.CVPR 2022 · 80 citations
