Towards Fine-Grained HBOE with Rendered Orientation Set and Laplace Smoothing
Ruisi Zhao, Mingming Li, Zheng Yang, Binbin Lin, Xiaohui Zhong, Xiaobo Ren, Deng Cai, Boxi Wu
Abstract
Human body orientation estimation (HBOE) aims to estimate the orientation of a human body relative to the camera’s frontal view. Despite recent advancements in this field, there still exist limitations in achieving fine-grained results. We identify certain defects and propose corresponding approaches as follows: 1). Existing datasets suffer from non-uniform angle distributions, resulting in sparse image data for certain angles. To provide comprehensive and high-quality data, we introduce RMOS (Rendered Model Orientation Set), a rendered dataset comprising 150K accurately labeled human instances with a wide range of orientations. 2). Directly using one-hot vector as labels may overlook the similarity between angle labels, leading to poor supervision. And converting the predictions from radians to degrees enlarges the regression error. To enhance supervision, we employ Laplace smoothing to vectorize the label, which contains more information. For fine-grained predictions, we adopt weighted Smooth-L1-loss to align predictions with the smoothed-label, thus providing robust supervision. 3). Previous works ignore body-part-specific information, resulting in coarse predictions. By employing local-window self-attention, our model could utilize different body part information for more precise orientation estimations. We validate the effectiveness of our method in the benchmarks with extensive experiments and show that our method outperforms state-of-the-art. Project is available at: https://github.com/Whalesong-zrs/Towards-Fine-grained-HBOE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d005519-1f78-4bca-bd36-bba2509ad9aaCited by top-tier papers1
Ask how each one uses itBuilds on7
- CoAtNet: Marrying Convolution and Attention for All Data SizesZihang Dai, Hanxiao Liu, Quoc V. Le, Mingxing TanNeurIPS 2021 · 1,747 citations
- Early Convolutions Help Transformers See BetterTete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell et al.NeurIPS 2021 · 974 citations
- Local Relation Networks for Image RecognitionHan Hu, Zheng Zhang, Zhenda Xie, Stephen LinICCV 2019 · 555 citations
- HRFormer: High-Resolution Vision Transformer for Dense PredictYuhui Yuan, Rao Fu, Lang Huang, Weihong Lin et al.NeurIPS 2021 · 357 citations
- Generative Models as a Data Source for Multiview Representation LearningAli Jahanian, Xavier Puig, Yonglong Tian, Phillip IsolaICLR 2022 · 148 citations
Related papers
- MEBOW: Monocular Estimation of Body Orientation in the WildChenyan Wu, Yukun Chen, Jiajia Luo, Che-Chun Su et al.CVPR 2020
- BGHR: Bridging the Gap Between HBox-Supervised and RBox-Supervised Oriented Object Detection via Adaptive Fine-Grained Sample MiningChenlin Fu, Yingying ZhuAAAI 2025 · 2 citations
- xR-EgoPose: Egocentric 3D Human Pose From an HMD CameraDenis Tomè, Patrick Peluse, Lourdes Agapito, Hernán BadinoICCV 2019 · 140 citations
- Zolly: Zoom Focal Length Correctly for Perspective-Distorted Human Mesh ReconstructionWenjia Wang, Yongtao Ge, Haiyi Mei, Zhongang Cai et al.ICCV 2023 · 51 citations
- H2RBox: Horizontal Box Annotation is All You Need for Oriented Object DetectionXue Yang, Gefan Zhang, Wentong Li, Yue Zhou et al.ICLR 2023 · 24 citations
