Gesture-aware Interactive Machine Teaching with In-situ Object Annotations
Zhongyi Zhou, Koji Yatani
Abstract
Interactive Machine Teaching (IMT) systems allow non-experts to easily create Machine Learning (ML) models. However, existing vision-based IMT systems either ignore annotations on the objects of interest or require users to annotate in a post-hoc manner. Without the annotations on objects, the model may misinterpret the objects using unrelated features. Post-hoc annotations cause additional workload, which diminishes the usability of the overall model building process. In this paper, we develop LookHere, which integrates in-situ object annotations into vision-based IMT. LookHere exploits users’ deictic gestures to segment the objects of interest in real time. This segmentation information can be additionally used for training. To achieve the reliable performance of this object segmentation, we utilize our custom dataset called HuTics, including 2040 front-facing images of deictic gestures toward various objects by 170 people. The quantitative results of our user study showed that participants were 16.3 times faster in creating a model with our system compared to a standard IMT system with a post-hoc annotation process while demonstrating comparable accuracies. Additionally, models created by our system showed a significant accuracy improvement (ΔmIoU = 0.466) in segmenting the objects of interest compared to those without annotations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7ffd83f-2ea6-4b99-b01a-2b0df3b8c17cCited by top-tier papers8
- Teachable Reality: Prototyping Tangible Augmented Reality with Everyday Objects by Leveraging Interactive Machine TeachingKyzyl Monteiro, Ritik Vatsal, Neil Chulpongsatorn, Aman Parnami et al.CHI 2023 · 55 citations
- Vision-Based Multimodal Interfaces: A Survey and Taxonomy for Enhanced Context-Aware System DesignYongquan 'Owen' Hu, Jingyu Tang, Xinya Gong, Zhongyi Zhou et al.CHI 2025 · 37 citations
- ThingShare: Ad-Hoc Digital Copies of Physical Objects for Sharing Things in Video MeetingsErzhen Hu, Jens Emil Sloth Grønbæk, Wen Ying, Ruofei Du et al.CHI 2023 · 30 citations
- Mapping the Design Space of Teachable Social Media Feed ExperiencesK. J. Kevin Feng, Xander Koo, Lawrence Tan, Amy S. Bruckman et al.CHI 2024 · 20 citations
- InstructPipe: Generating Visual Blocks Pipelines with Human Instructions and LLMsZhongyi Zhou, Jing Jin, Vrushank Phadnis, Xiuxiu Yuan et al.CHI 2025 · 10 citations
Builds on12
- Reconstructing Hand-Object Interactions in the WildZhe Cao, Ilija Radosavovic, Angjoo Kanazawa, Jitendra MalikICCV 2021 · 184 citations
- No Explainability without Accountability: An Empirical Study of Explanations and Feedback in Interactive MLAlison Smith-Renner, Ron Fan, Melissa Birchfield, Tongshuang Wu et al.CHI 2020 · 117 citations
- Panoptic 3D Scene Reconstruction From a Single RGB ImageManuel Dahnert, Ji Hou, Matthias Nießner, Angela DaiNeurIPS 2021 · 106 citations
- Hand Interfaces: Using Hands to Imitate Objects in AR/VR for Expressive InteractionsSiyou Pei, Alexander Chen, Jaewook Lee, Yang ZhangCHI 2022 · 93 citations
- Marcelle: Composing Interactive Machine Learning Workflows and InterfacesJules Françoise, Baptiste Caramiaux, Téo SanchezUIST 2021 · 37 citations
Related papers
- LocTex: Learning Data-Efficient Visual Representations from Localized Textual SupervisionZhijian Liu, Simon Stent, Jie Li, John Gideon et al.ICCV 2021 · 10 citations
- Overthere: A Simple and Intuitive Object Registration Method for an Absolute Mid-air Pointing InterfaceHyunggoog Seo, Jaedong Kim, Kwanggyoon Seo, Bumki Kim et al.UbiComp 2021 · 5 citations
- YouRefIt: Embodied Reference Understanding with Language and GestureYixin Chen, Qing Li, Deqian Kong, Yik Lun Kei et al.ICCV 2021 · 57 citations
- MultiHGR: Multi-Task Hand Gesture Recognition with Cross-Modal Wrist-Worn DevicesMengxia Lyu, Hao Zhou, Kaiwen Guo, Wangqiu Zhou et al.INFOCOM 2024 · 4 citations
- ViSpeak: Visual Instruction Feedback in Streaming VideosShenghao Fu, Qize Yang, Yuan-Ming Li, Yi-Xing Peng et al.ICCV 2025 · 42 citations
