Semantic Human Parsing via Scalable Semantic Transfer Over Multiple Label Domains
Jie Yang, Chaoqun Wang, Zhen Li, Junle Wang, Ruimao Zhang
Abstract
This paper presents Scalable Semantic Transfer (SST), a novel training paradigm, to explore how to leverage the mutual benefits of the data from different label domains (i.e. various levels of label granularity) to train a powerful human parsing network. In practice, two common application scenarios are addressed, termed universal parsing and dedicated parsing, where the former aims to learn homogeneous human representations from multiple label domains and switch predictions by only using different segmentation heads, and the latter aims to learn a specific domain prediction while distilling the semantic knowledge from other domains. The proposed SST has the following appealing benefits: (1) it can capably serve as an effective training scheme to embed semantic associations of human body parts from multiple label domains into the human representation learning process; (2) it is an extensible semantic transfer framework without predetermining the overall relations of multiple label domains, which allows continuously adding human parsing datasets to promote the training. (3) the relevant modules are only used for auxiliary training and can be removed during inference, eliminating the extra reasoning cost. Experimental results demonstrate SST can effectively achieve promising universal human parsing performance as well as impressive improvements compared to its counterparts on three human parsing benchmarks (i.e., PASCAL-Person-Part, ATR, and CIHP). Code is available at https://github.com/yangjie-cv/SST .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7a6694fe-86df-409d-afbe-8f0464b602b7Cited by top-tier papers2
- HumanMAC: Masked Motion Completion for Human Motion PredictionLing-Hao Chen, Jiawei Zhang, Yewen Li, Yiren Pang et al.ICCV 2023 · 106 citations
- KptLLM: Unveiling the Power of Large Language Model for Keypoint ComprehensionJie Yang, Wang Zeng, Sheng Jin, Lumin Xu et al.NeurIPS 2024 · 9 citations
Builds on17
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- LAVT: Language-Aware Vision Transformer for Referring Image SegmentationZhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen et al.CVPR 2022 · 319 citations
- Deep Hierarchical Semantic SegmentationLiulei Li, Tianfei Zhou, Wenguan Wang, Jianwu Li et al.CVPR 2022 · 181 citations
Related papers
- Grapy-ML: Graph Pyramid Mutual Learning for Cross-Dataset Human ParsingHaoyu He, Jing Zhang, Qiming Zhang, Dacheng TaoAAAI 2020 · 65 citations
- Part-Aware Context Network for Human ParsingXiaomei Zhang, Yingying Chen, Bingke Zhu, Jinqiao Wang et al.CVPR 2020
- Universal-RCNN: Universal Object Detector via Transferable Graph R-CNNHang Xu, Linpu Fang, Xiaodan Liang, Wenxiong Kang et al.AAAI 2020 · 26 citations
- DaTaSeg: Taming a Universal Multi-Dataset Multi-Task Segmentation ModelXiuye Gu, Yin Cui, Jonathan Huang, Abdullah Rashwan et al.NeurIPS 2023 · 40 citations
- CDGNet: Class Distribution Guided Network for Human ParsingKunliang Liu, Ouk Choi, Jianming Wang, Wonjun HwangCVPR 2022 · 44 citations
