Open-Category Human-Object Interaction Pre-training via Language Modeling Framework
Sipeng Zheng, Boshen Xu, Qin Jin
Abstract
Human-object interaction (HOI) has long been plagued by the conflict between limited supervised data and a vast number of possible interaction combinations in real life. Current methods trained from closed-set data predict HOIs as fixed-dimension logits, which restricts their scalability to open-set categories. To address this issue, we introduce OpenCat, a language modeling framework that reformulates HOI prediction as sequence generation. By converting HOI triplets into a token sequence through a serialization scheme, our model is able to exploit the open-set vocabulary of the language modeling framework to predict novel interaction classes with a high degree of freedom. In addition, inspired by the great success of visionlanguage pre-training, we collect a large amount of weaklysupervised data related to HOI from image-caption pairs, and devise several auxiliary proxy tasks, including soft relational matching and human-object relation prediction, to pre-train our model. Extensive experiments show that our OpenCat significantly boosts HOI performance, particularly on a broad range of rare and unseen categories.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e88066fc-25ae-4e99-a54c-fa0e1d6ce7c2Cited by top-tier papers13
- Vision-Language-Action Pretraining from Large-Scale Human VideosHao Luo, Yicheng Feng, Wanpeng Zhang, Sipeng Zheng et al.ICML 2026 · 104 citations
- Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion ModelsLiulei Li, Wenguan Wang, Yi YangNeurIPS 2024 · 29 citations
- Toward Open-Set Human Object Interaction DetectionMingrui Wu, Yuqi Liu, Jiayi Ji, Xiaoshuai Sun et al.AAAI 2024 · 12 citations
- Discovering Syntactic Interaction Clues for Human-Object Interaction DetectionJinguo Luo, Weihong Ren, Weibo Jiang, Xi'ai Chen et al.CVPR 2024 · 10 citations
- Open-Vocabulary Hoi Detection With Interaction-Aware Prompt and Concept CalibrationTing Lei, Shaofeng Yin, Qingchao Chen, Yuxin Peng et al.ICCV 2025 · 6 citations
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Unifying Vision-and-Language Tasks via Text GenerationJaemin Cho, Jie Lei, Hao Tan, Mohit BansalICML 2021 · 624 citations
- Pix2seq: A Language Modeling Framework for Object DetectionTing Chen, Saurabh Saxena, Lala Li, David J. Fleet et al.ICLR 2022 · 435 citations
Related papers
- Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation ModelsYichao Cao, Qingfei Tang, Xiu Su, Song Chen et al.NeurIPS 2023 · 64 citations
- Towards Open-vocabulary HOI Detection with Calibrated Vision-language Models and Locality-aware QueriesZhenhao Yang, Xin Liu, Deqiang Ouyang, Guiduo Duan et al.ACM MM 2024 · 5 citations
- HOICLIP: Efficient Knowledge Transfer for HOI Detection with Vision-Language ModelsShan Ning, Longtian Qiu, Yongfei Liu, Xuming HeCVPR 2023
- Detecting Human-Object Interaction via Fabricated Compositional LearningZhi Hou, Baosheng Yu, Yu Qiao, Xiaojiang Peng et al.CVPR 2021
- RLIP: Relational Language-Image Pre-training for Human-Object Interaction DetectionHangjie Yuan, Jianwen Jiang, Samuel Albanie, Tao Feng et al.NeurIPS 2022 · 88 citations
