Seeing Eye to Eye: Enabling Cognitive Alignment Through Shared First-Person Perspective in Human-AI Collaboration: Seeing Eye to Eye
Zhuyu Teng, Pei Chen, Yichen Cai, Ruoqing Lu, Zhaoqu Jiang, Jiayang Li, Weitao You, Lingyun Sun
摘要
Despite advances in multimodal AI, current vision-based assistants often remain inefficient in collaborative tasks. We identify two key gulfs: a communication gulf, where users must translate rich parallel intentions into verbal commands due to the channel mismatch, and an understanding gulf, where AI struggles to interpret subtle embodied cues. To address these, we propose Eye2Eye, a framework that leverages first-person perspective as a channel for human-AI cognitive alignment. It integrates three components: (1) joint attention coordination for fluid focus alignment, (2) revisable memory to maintain evolving common ground, and (3) reflective feedback allowing users to clarify and refine AI’s understanding. We implement this framework in an AR prototype and evaluate it through a user study and a post-hoc pipeline evaluation. Results show that Eye2Eye significantly reduces task completion time and interaction load while increasing trust, demonstrating its components work in concert to improve collaboration.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Understanding Design Collaboration Between Designers and Artificial Intelligence: A Systematic Literature ReviewYang Shi, Tian Gao, Xiaohan Jiao, Nan CaoCSCW 2023 · 被引用 170 次
- AdapTutAR: An Adaptive Tutoring System for Machine Tasks in Augmented RealityGaoping Huang, Xun Qian, Tianyi Wang, Fagun Patel 等CHI 2021 · 被引用 93 次
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented RealityJaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu 等CHI 2024 · 被引用 86 次
- Augmented Reality for Older Adults: Exploring Acceptability of Virtual Coaches for Home-based Balance Training in an Aging PopulationFariba Mostajeran, Frank Steinicke, Oscar Javier Ariza Nunez, Dimitrios A. Gatsios 等CHI 2020 · 被引用 80 次
- Satori 悟り: Towards Proactive AR Assistant with Belief-Desire-Intention User ModelingChenyi Li, Guande Wu, Gromit Yeuk-Yin Chan, Dishita G. Turakhia 等CHI 2025 · 被引用 49 次
相关 Paper
- Persistent Assistant: Seamless Everyday AI Interactions via Intent Grounding and Multimodal FeedbackHyunsung Cho, Jacqui Fashimpaur, Naveen Sendhilnathan, Jonathan Browder 等CHI 2025 · 被引用 10 次
- Supporting Joint Attention and Body Alignment Between Remote Users Walking OutdoorsTakeru Yazaki, Noriyasu Obushi, Kuniharu Sakurada, Hideaki Kuzuoka 等CSCW 2026
- Understanding User Reliance on AI in Assisted Decision-MakingShiye Cao, Chien-Ming HuangCSCW 2022 · 被引用 68 次
- Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent DisambiguationSicheng Yang, Yukai Huang, Weitong Cai, Shitong Sun 等AAAI 2026
- Ego2Web: A Web Agent Benchmark Grounded in Egocentric VideosShoubin Yu, Lei Shu, Antoine Yang, Yao Fu 等CVPR 2026 · 被引用 4 次
