Understanding Embodied Reference with Touch-Line Transformer
Yang Li, Xiaoxue Chen, Hao Zhao, Jiangtao Gong, Guyue Zhou, Federico Rossano, Yixin Zhu
Abstract
We study embodied reference understanding, the task of locating referents using embodied gestural signals and language references. Human studies have revealed that objects referred to or pointed to do not lie on the elbow-wrist line, a common misconception; instead, they lie on the so-called virtual touch line. However, existing human pose representations fail to incorporate the virtual touch line. To tackle this problem, we devise the touch-line transformer: It takes as input tokenized visual and textual features and simultaneously predicts the referent's bounding box and a touch-line vector. Leveraging this touch-line prior, we further devise a geometric consistency loss that encourages the co-linearity between referents and touch lines. Using the touch-line as gestural information improves model performances significantly. Experiments on the YouRefIt dataset show our method achieves a +25.0% accuracy improvement under the 0.75 IoU criterion, closing 63.6% of the gap between model and human performances. Furthermore, we computationally verify prior human studies by showing that computational models more accurately locate referents when using the virtual touch line than when using the elbow-wrist line.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- HUMANISE: Language-conditioned Human Motion Generation in 3D ScenesZan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu et al.NeurIPS 2022 · 207 citations
- RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression SegmentationChangli Wu, Qi Chen, Jiayi Ji, Haowei Wang et al.NeurIPS 2024 · 16 citations
- ScanERU: Interactive 3D Visual Grounding Based on Embodied Reference UnderstandingZiyang Lu, Yunqiang Pei, Guoqing Wang, Peiwei Li et al.AAAI 2024 · 12 citations
- TeleAware Robot: Designing Awareness-augmented Telepresence Robot for Remote Collaborative LocomotionRuyi Li, Yaxin Zhu, Min Liu, Yihang Zeng et al.UbiComp 2024 · 7 citations
- RESAnything: Attribute Prompting for Arbitrary Referring SegmentationRuiqi Wang, Hao ZhangNeurIPS 2025 · 6 citations
Builds on10
- Attention Augmented Convolutional NetworksIrwan Bello, Barret Zoph, Quoc Le, Ashish Vaswani et al.ICCV 2019 · 1,149 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- Gaze360: Physically Unconstrained Gaze Estimation in the WildPetr Kellnhofer, Adrià Recasens, Simon Stent, Wojciech Matusik et al.ICCV 2019 · 469 citations
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang et al.ICCV 2019 · 441 citations
- MAT: Mask-Aware Transformer for Large Hole Image InpaintingWenbo Li, Zhe Lin, Kun Zhou, Lu Qi et al.CVPR 2022 · 382 citations
Related papers
- YouRefIt: Embodied Reference Understanding with Language and GestureYixin Chen, Qing Li, Deqian Kong, Yik Lun Kei et al.ICCV 2021 · 57 citations
- Ges3ViG : Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference UnderstandingAtharv Mahesh Mane, Dulanga Weerakoon, Vigneshwaran Subbaraju, Sougata Sen et al.CVPR 2025
- Mono3DVG: 3D Visual Grounding in Monocular ImagesYang Zhan, Yuan Yuan, Zhitong XiongAAAI 2024 · 38 citations
- Toward Human Deictic Gesture Target EstimationXu Cao, Pranav Virupaksha, Sangmin Lee, Bolin Lai et al.NeurIPS 2025 · 3 citations
- InteractVLM: 3D Interaction Reasoning from 2D Foundational ModelsSai Kumar Dwivedi, Dimitrije Antic, Shashank Tripathi, Omid Taheri et al.CVPR 2025
