PaStaNet: Toward Human Activity Knowledge Engine
Yong-Lu Li, Liang Xu, Xinpeng Liu, Xijie Huang, Yue Xu, Shiyi Wang, Haoshu Fang, Ze Ma, Mingyang Chen, Cewu Lu
摘要
Existing image-based activity understanding methods mainly adopt direct mapping, i.e. from image to activity concepts, which may encounter performance bottleneck since the huge gap. In light of this, we propose a new path: infer human part states first and then reason out the activities based on part-level semantics. Human Body Part States (PaSta) are fine-grained action semantic tokens, e.g. hand, hold, something , which can compose the activities and help us step toward human activity knowledge engine. To fully utilize the power of PaSta, we build a largescale knowledge base PaStaNet, which contains 7M+ PaSta annotations. And two corresponding models are proposed: first, we design a model named Activity2Vec to extract PaSta features, which aim to be general representations for various activities. Second, we use a PaSta-based Reasoning method to infer activities. Promoted by PaStaNet, our method achieves significant improvements, e.g. 6.4 and 13.9 mAP on full and one-shot sets of HICO in supervised learning, and 3.2 and 4.2 mAP on V-COCO and images-based AVA in transfer learning. Code and data are available at http://hake-mvig.cn/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper54
- HOI Analysis: Integrating and Decomposing Human-Object InteractionYong-Lu Li, Xinpeng Liu, Xiaoqian Wu, Yizhuo Li 等NeurIPS 2020 · 被引用 152 次
- Three Steps to Multimodal Trajectory Prediction: Modality Clustering, Classification and SynthesisJianhua Sun, Yuxuan Li, Haoshu Fang, Cewu LuICCV 2021 · 被引用 91 次
- MSTR: Multi-Scale Transformer for End-to-End Human-Object Interaction DetectionBumsoo Kim, Jonghwan Mun, Kyoung-Woon On, Minchul Shin 等CVPR 2022 · 被引用 80 次
- Learning Transferable Human-Object Interaction Detector with Natural Language SupervisionSuchen Wang, Yueqi Duan, Henghui Ding, Yap-Peng Tan 等CVPR 2022 · 被引用 66 次
- DIRV: Dense Interaction Region Voting for End-to-End Human-Object Interaction DetectionHaoshu Fang, Yichen Xie, Dian Shao, Cewu LuAAAI 2021 · 被引用 66 次
它引用的顶会 Paper2
相关 Paper
- Visual Knowledge Graph for Human Action Reasoning in VideosYue Ma, Yali Wang, Yue Wu, Ziyu Lyu 等ACM MM 2022 · 被引用 29 次
- ART: rule bAsed futuRe-inference deducTionMengze Li, Tianqi Zhao, Jionghao Bai, Baoyi He 等EMNLP 2023 · 被引用 2 次
- Hierarchical Human Parsing With Typed Part-Relation ReasoningWenguan Wang, Hailong Zhu, Jifeng Dai, Yanwei Pang 等CVPR 2020
- Semantic Human Parsing via Scalable Semantic Transfer Over Multiple Label DomainsJie Yang, Chaoqun Wang, Zhen Li, Junle Wang 等CVPR 2023
- PASTA: Table-Operations Aware Fact Verification via Sentence-Table Cloze Pre-trainingZihui Gu, Ju Fan, Nan Tang, Preslav Nakov 等EMNLP 2022 · 被引用 19 次
