Interaction-aware Joint Attention Estimation Using People Attributes
Chihiro Nakatani, Hiroaki Kawashima, Norimichi Ukita
摘要
This paper proposes joint attention estimation in a single image. Different from related work in which only the gazerelated attributes of people are independently employed, (i) their locations and actions are also employed as contextual cues for weighting their attributes, and (ii) interactions among all of these attributes are explicitly modeled in our method. For the interaction modeling, we propose a novel Transformer-based attention network to encode joint attention as low-dimensional features. We introduce a specialized MLP head with positional embedding to the Transformer so that it predicts pixelwise confidence of joint attention for generating the confidence heatmap. This pixelwise prediction improves the heatmap accuracy by avoiding the ill-posed problem in which the high-dimensional heatmap is predicted from the low-dimensional features. The estimated joint attention is further improved by being integrated with general image-based attention estimation. Our method outperforms SOTA methods quantitatively in comparative experiments. Code: https://github.com/chihina/ PJAE .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Learning Group Activity Features Through Person Attribute PredictionChihiro Nakatani, Hiroaki Kawashima, Norimichi UkitaCVPR 2024 · 被引用 4 次
- Gaze-LLE: Gaze Target Estimation via Large-Scale Learned EncodersFiona Ryan, Ajay Bati, Sangmin Lee, Daniel Bolya 等CVPR 2025
- Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned RepresentationsSangmin Lee, Bolin Lai, Fiona Ryan, Bikram Boote 等CVPR 2024
它引用的顶会 Paper16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 被引用 672 次
- Gaze360: Physically Unconstrained Gaze Estimation in the WildPetr Kellnhofer, Adrià Recasens, Simon Stent, Wojciech Matusik 等ICCV 2019 · 被引用 469 次
- Joint Learning of Saliency Detection and Weakly Supervised Semantic SegmentationYu Zeng, Yun-Zhi Zhuge, Huchuan Lu, Lihe ZhangICCV 2019 · 被引用 190 次
相关 Paper
- A2J-Transformer: Anchor-to-Joint Transformer Network for 3D Interacting Hand Pose Estimation from a Single RGB ImageChanglong Jiang, Yang Xiao, Cunlin Wu, Mingyang Zhang 等CVPR 2023
- TransGOP: Transformer-Based Gaze Object PredictionBinglu Wang, Chenxi Guo, Yang Jin, Haisheng Xia 等AAAI 2024 · 被引用 8 次
- Object-aware Gaze Target DetectionFrancesco Tonini, Nicola Dall'Asen, Cigdem Beyan, Elisa RicciICCV 2023 · 被引用 38 次
- QPIC: Query-Based Pairwise Human-Object Interaction Detection With Image-Wide Contextual InformationMasato Tamura, Hiroki Ohashi, Tomoaki YoshinagaCVPR 2021
- Keypoint Transformer: Solving Joint Identification in Challenging Hands and Object Interactions for Accurate 3D Pose EstimationShreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, Vincent LepetitCVPR 2022 · 被引用 155 次
