From Isolated Islands to Pangea: Unifying Semantic Space for Human Action Understanding
Yong-Lu Li, Xiaoqian Wu, Xinpeng Liu, Zehao Wang, Yiming Dou, Yikun Ji, Junyi Zhang, Yixing Li, Xudong Lu, Jingru Tan, Cewu Lu
Abstract
Action understanding has attracted long-term attention. It can be formed as the mapping from the physical space to the semantic space. Typically, researchers built datasets according to idiosyncratic choices to define classes and push the envelope of benchmarks respectively. Datasets are incompatible with each other like "Isolated Islands" due to semantic gaps and various class granularities, e.g., do housework in dataset A and wash plate in dataset B. We argue that we need a more principled semantic space to concentrate the community efforts and use all datasets together to pursue generalizable action learning. To this end, we design a structured action semantic space given verb taxonomy hierarchy and covering massive actions. By aligning the classes of previous datasets to our semantic space, we gather (image/video/skeleton/MoCap) datasets into a unified database in a unified label system, i.e., bridging "isolated islands" into a "Pangea". Accordingly, we propose a novel model mapping from the physical space to semantic space to fully use Pangea. In extensive experiments, our new system shows significant superiority, especially in transfer learning. Our code and data will be made public at https://mvig-rhos.com/pangea.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5e5e73eb-1e96-43f3-9ddb-d29cd0dd3f29Cited by top-tier papers15
- WorDepth: Variational Language Prior for Monocular Depth EstimationZiyao Zeng, Daniel Wang, Fengyu Yang, Hyoungseob Park et al.CVPR 2024 · 20 citations
- Beyond Object Recognition: A New Benchmark towards Object Concept LearningYonglu Li, Yue Xu, Xinyu Xu, Xiaohan Mao et al.ICCV 2023 · 12 citations
- KinMo: Kinematic-Aware Human Motion Understanding and GenerationPengfei Zhang, Pinxin Liu, Pablo Garrido, Hyeongwoo Kim et al.ICCV 2025 · 9 citations
- Generating Attribute-Aware Human Motions from Textual PromptXinghan Wang, Kun Xu, Fei Li, Cao Sheng et al.AAAI 2026
- Reconstructing In-the-Wild Open-Vocabulary Human-Object InteractionsBoran Wen, Dingbang Huang, Zichen Zhang, Jiahong Zhou et al.CVPR 2025
Builds on34
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
Related papers
- BABEL: Bodies, Action and Behavior With English LabelsAbhinanda R. Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou, Alejandra Quiros-Ramirez et al.CVPR 2021
- Simple Multi-dataset DetectionXingyi Zhou, Vladlen Koltun, Philipp KrähenbühlCVPR 2022 · 85 citations
- A Large-scale Study of Spatiotemporal Representation Learning with a New Benchmark on Action RecognitionAndong Deng, Taojiannan Yang, Chen ChenICCV 2023 · 18 citations
- Generative Action Description Prompts for Skeleton-based Action RecognitionWangmeng Xiang, Chao Li, Yuxuan Zhou, Biao Wang et al.ICCV 2023 · 84 citations
- Zero-Shot Learning for IMU-Based Activity Recognition Using Video EmbeddingsCatherine Tong, Jinchen Ge, Nicholas D. LaneUbiComp 2022 · 39 citations
