Scene-aware Human Pose Generation using Transformer
Jieteng Yao, Junjie Chen, Li Niu, Bin Sheng
Abstract
Affordance learning considers the interaction opportunities for an actor in the scene and thus has wide application in scene understanding and intelligent robotics. In this paper, we focus on contextual affordance learning, i.e., using affordance as context to generate a reasonable human pose in a scene. Existing scene-aware human pose generation methods could be divided into two categories depending on whether using pose templates. Our proposed method belongs to the template-based category, which benefits from the representative pose templates. Moreover, inspired by recent transformer-based methods, we associate each query embedding with a pose template, and use the interaction between query embeddings and scene feature map to effectively predict the scale and offsets for each pose template. In addition, we employ knowledge distillation to facilitate the offset learning given the predicted scale. Comprehensive experiments on Sitcom dataset demonstrate the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- HuMoR: 3D Human Motion Model for Robust Pose EstimationDavis Rempe, Tolga Birdal, Aaron Hertzmann, Jimei Yang et al.ICCV 2021 · 398 citations
- AvatarCLIP: zero-shot text-driven generation and animation of 3D avatarsFangzhou Hong, Mingyuan Zhang, Liang Pan, Zhongang Cai et al.SIGGRAPH 2022 · 213 citations
- Pose Recognition With Cascade TransformersKe Li, Shijie Wang, Xiang Zhang, Yifan Xu et al.CVPR 2021
- UniPose: Unified Human Pose Estimation in Single Images and VideosBruno Artacho, Andreas E. SavakisCVPR 2020
Related papers
- Understanding 3D Object Interaction from a Single ImageShengyi Qian, David F. FouheyICCV 2023 · 35 citations
- Affordance Grounding from Demonstration Video to Target ImageJoya Chen, Difei Gao, Kevin Qinghong Lin, Mike Zheng ShouCVPR 2023
- Putting People in Their Place: Affordance-Aware Human Insertion into ScenesSumith Kulal, Tim Brooks, Alex Aiken, Jiajun Wu et al.CVPR 2023
- Order-aware Human Interaction ManipulationMandi Luo, Jie Cao, Ran HeACM MM 2022 · 1 citation
- VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic ManipulationHanzhi Chen, Boyang Sun, Anran Zhang, Marc Pollefeys et al.CVPR 2025
