Gesture-aware Interactive Machine Teaching with In-situ Object Annotations
Zhongyi Zhou, Koji Yatani
摘要
Interactive Machine Teaching (IMT) systems allow non-experts to easily create Machine Learning (ML) models. However, existing vision-based IMT systems either ignore annotations on the objects of interest or require users to annotate in a post-hoc manner. Without the annotations on objects, the model may misinterpret the objects using unrelated features. Post-hoc annotations cause additional workload, which diminishes the usability of the overall model building process. In this paper, we develop LookHere, which integrates in-situ object annotations into vision-based IMT. LookHere exploits users’ deictic gestures to segment the objects of interest in real time. This segmentation information can be additionally used for training. To achieve the reliable performance of this object segmentation, we utilize our custom dataset called HuTics, including 2040 front-facing images of deictic gestures toward various objects by 170 people. The quantitative results of our user study showed that participants were 16.3 times faster in creating a model with our system compared to a standard IMT system with a post-hoc annotation process while demonstrating comparable accuracies. Additionally, models created by our system showed a significant accuracy improvement (ΔmIoU = 0.466) in segmenting the objects of interest compared to those without annotations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Teachable Reality: Prototyping Tangible Augmented Reality with Everyday Objects by Leveraging Interactive Machine TeachingKyzyl Monteiro, Ritik Vatsal, Neil Chulpongsatorn, Aman Parnami 等CHI 2023 · 被引用 55 次
- Vision-Based Multimodal Interfaces: A Survey and Taxonomy for Enhanced Context-Aware System DesignYongquan 'Owen' Hu, Jingyu Tang, Xinya Gong, Zhongyi Zhou 等CHI 2025 · 被引用 37 次
- ThingShare: Ad-Hoc Digital Copies of Physical Objects for Sharing Things in Video MeetingsErzhen Hu, Jens Emil Sloth Grønbæk, Wen Ying, Ruofei Du 等CHI 2023 · 被引用 30 次
- Mapping the Design Space of Teachable Social Media Feed ExperiencesK. J. Kevin Feng, Xander Koo, Lawrence Tan, Amy S. Bruckman 等CHI 2024 · 被引用 20 次
- InstructPipe: Generating Visual Blocks Pipelines with Human Instructions and LLMsZhongyi Zhou, Jing Jin, Vrushank Phadnis, Xiuxiu Yuan 等CHI 2025 · 被引用 10 次
它引用的顶会 Paper12
- Reconstructing Hand-Object Interactions in the WildZhe Cao, Ilija Radosavovic, Angjoo Kanazawa, Jitendra MalikICCV 2021 · 被引用 184 次
- No Explainability without Accountability: An Empirical Study of Explanations and Feedback in Interactive MLAlison Smith-Renner, Ron Fan, Melissa Birchfield, Tongshuang Wu 等CHI 2020 · 被引用 117 次
- Panoptic 3D Scene Reconstruction From a Single RGB ImageManuel Dahnert, Ji Hou, Matthias Nießner, Angela DaiNeurIPS 2021 · 被引用 106 次
- Hand Interfaces: Using Hands to Imitate Objects in AR/VR for Expressive InteractionsSiyou Pei, Alexander Chen, Jaewook Lee, Yang ZhangCHI 2022 · 被引用 93 次
- Marcelle: Composing Interactive Machine Learning Workflows and InterfacesJules Françoise, Baptiste Caramiaux, Téo SanchezUIST 2021 · 被引用 37 次
相关 Paper
- LocTex: Learning Data-Efficient Visual Representations from Localized Textual SupervisionZhijian Liu, Simon Stent, Jie Li, John Gideon 等ICCV 2021 · 被引用 10 次
- Overthere: A Simple and Intuitive Object Registration Method for an Absolute Mid-air Pointing InterfaceHyunggoog Seo, Jaedong Kim, Kwanggyoon Seo, Bumki Kim 等UbiComp 2021 · 被引用 5 次
- YouRefIt: Embodied Reference Understanding with Language and GestureYixin Chen, Qing Li, Deqian Kong, Yik Lun Kei 等ICCV 2021 · 被引用 57 次
- MultiHGR: Multi-Task Hand Gesture Recognition with Cross-Modal Wrist-Worn DevicesMengxia Lyu, Hao Zhou, Kaiwen Guo, Wangqiu Zhou 等INFOCOM 2024 · 被引用 4 次
- ViSpeak: Visual Instruction Feedback in Streaming VideosShenghao Fu, Qize Yang, Yuan-Ming Li, Yi-Xing Peng 等ICCV 2025 · 被引用 42 次
