Seeing Eye to Eye: Enabling Cognitive Alignment Through Shared First-Person Perspective in Human-AI Collaboration: Seeing Eye to Eye
Zhuyu Teng, Pei Chen, Yichen Cai, Ruoqing Lu, Zhaoqu Jiang, Jiayang Li, Weitao You, Lingyun Sun
Abstract
Despite advances in multimodal AI, current vision-based assistants often remain inefficient in collaborative tasks. We identify two key gulfs: a communication gulf, where users must translate rich parallel intentions into verbal commands due to the channel mismatch, and an understanding gulf, where AI struggles to interpret subtle embodied cues. To address these, we propose Eye2Eye, a framework that leverages first-person perspective as a channel for human-AI cognitive alignment. It integrates three components: (1) joint attention coordination for fluid focus alignment, (2) revisable memory to maintain evolving common ground, and (3) reflective feedback allowing users to clarify and refine AI’s understanding. We implement this framework in an AR prototype and evaluate it through a user study and a post-hoc pipeline evaluation. Results show that Eye2Eye significantly reduces task completion time and interaction load while increasing trust, demonstrating its components work in concert to improve collaboration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dbf19e1a-3cee-4967-8c5e-5339107f3df9Builds on22
- Understanding Design Collaboration Between Designers and Artificial Intelligence: A Systematic Literature ReviewYang Shi, Tian Gao, Xiaohan Jiao, Nan CaoCSCW 2023 · 170 citations
- AdapTutAR: An Adaptive Tutoring System for Machine Tasks in Augmented RealityGaoping Huang, Xun Qian, Tianyi Wang, Fagun Patel et al.CHI 2021 · 93 citations
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented RealityJaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu et al.CHI 2024 · 86 citations
- Augmented Reality for Older Adults: Exploring Acceptability of Virtual Coaches for Home-based Balance Training in an Aging PopulationFariba Mostajeran, Frank Steinicke, Oscar Javier Ariza Nunez, Dimitrios A. Gatsios et al.CHI 2020 · 80 citations
- Satori 悟り: Towards Proactive AR Assistant with Belief-Desire-Intention User ModelingChenyi Li, Guande Wu, Gromit Yeuk-Yin Chan, Dishita G. Turakhia et al.CHI 2025 · 49 citations
Related papers
- Persistent Assistant: Seamless Everyday AI Interactions via Intent Grounding and Multimodal FeedbackHyunsung Cho, Jacqui Fashimpaur, Naveen Sendhilnathan, Jonathan Browder et al.CHI 2025 · 10 citations
- Supporting Joint Attention and Body Alignment Between Remote Users Walking OutdoorsTakeru Yazaki, Noriyasu Obushi, Kuniharu Sakurada, Hideaki Kuzuoka et al.CSCW 2026
- Understanding User Reliance on AI in Assisted Decision-MakingShiye Cao, Chien-Ming HuangCSCW 2022 · 68 citations
- Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent DisambiguationSicheng Yang, Yukai Huang, Weitong Cai, Shitong Sun et al.AAAI 2026
- Ego2Web: A Web Agent Benchmark Grounded in Egocentric VideosShoubin Yu, Lei Shu, Antoine Yang, Yao Fu et al.CVPR 2026 · 4 citations
