Unified Pose Sequence Modeling
Lin Geng Foo, Tianjiao Li, Hossein Rahmani, Qiuhong Ke, Jun Liu
摘要
We propose a Unified Pose Sequence Modeling approach to unify heterogeneous human behavior understanding tasks based on pose data, e.g., action recognition, 3D pose estimation and 3D early action prediction. A major obstacle is that different pose-based tasks require different output data formats. Specifically, the action recognition and prediction tasks require class predictions as outputs, while 3D pose estimation requires a human pose output, which limits existing methods to leverage task-specific network architectures for each task. Hence, in this paper, we propose a novel Unified Pose Sequence (UPS) model to unify heterogeneous output formats for the aforementioned tasks by considering text-based action labels and coordinate-based human poses as language sequences. Then, by optimizing a single auto-regressive transformer, we can obtain a unified output sequence that can handle all the aforementioned tasks. Moreover, to avoid the interference brought by the heterogeneity between different tasks, a dynamic routing mechanism is also proposed to empower our UPS with the ability to learn which subsets of parameters should be shared among different tasks. To evaluate the efficacy of the proposed UPS, extensive experiments are conducted on four different tasks with four popular behavior understanding benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- LLMs are Good Action RecognizersHaoxuan Qu, Yujun Cai, Jun LiuCVPR 2024 · 被引用 37 次
- TCPFormer: Learning Temporal Correlation with Implicit Pose Proxy for 3D Human Pose EstimationJiajie Liu, Mengyuan Liu, Hong Liu, Wenhao LiAAAI 2025 · 被引用 27 次
- avaTTAR: Table Tennis Stroke Training with Embodied and Detached Visualization in Augmented RealityDizhi Ma, Xiyun Hu, Jingyu Shi, Mayank Patel 等UIST 2024 · 被引用 21 次
- Inter-X: Towards Versatile Human-Human Interaction AnalysisLiang Xu, Xintao Lv, Yichao Yan, Xin Jin 等CVPR 2024 · 被引用 18 次
- Skeleton-in-Context: Unified Skeleton Sequence Modeling with In-Context LearningXinshun Wang, Zhongbin Fang, Xia Li, Xiangtai Li 等CVPR 2024 · 被引用 12 次
它引用的顶会 Paper23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin 等ICML 2022 · 被引用 1,058 次
相关 Paper
- UniT: Multimodal Multitask Learning with a Unified TransformerRonghang Hu, Amanpreet SinghICCV 2021 · 被引用 354 次
- MotionBERT: A Unified Perspective on Learning Human Motion RepresentationsWentao Zhu, Xiaoxuan Ma, Zhaoyang Liu, Libin Liu 等ICCV 2023 · 被引用 322 次
- UniMotion: A Unified Motion Framework for Simulation, Prediction and PlanningNan Song, Junzhe Jiang, Jingyu Li, Xiatian Zhu 等NeurIPS 2025 · 被引用 2 次
- OmniViD: A Generative Framework for Universal Video UnderstandingJunke Wang, Dongdong Chen, Chong Luo, Bo He 等CVPR 2024 · 被引用 18 次
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang 等ICCV 2021 · 被引用 648 次
