GazeCoT: Unleashing Social Intelligence in Multimodal LLMs With Gaze-Informed Chain-of-Thought Reasoning
Zhoutong Ye, Xutong Wang, Chengwen Zhang, Ruiwen Zhang, Mingze Sun, Qinwei Li, Chun Yu, Yuanchun Shi
摘要
Social intelligence is vital for effective human-AI interaction. While LLMs demonstrate strong text-based social intelligence, the vision modality remains challenging due to the presence of non-verbal social cues. For example, gaze is the primary conveyor of social attention, yet it cannot be accurately perceived and understood by multimodal LLMs (MLLMs). Therefore, we propose GazeCoT, a pipeline using gaze estimation models to provide MLLMs with the attention of people in images or videos. The gaze information is provided as visual and text prompts compiled into a structured context to support MLLM social reasoning. Benchmark evaluation confirms that GazeCoT enhances MLLMs’ social intelligence by improving gaze perception. A user study in a challenging application involving parent-child interactions demonstrates that GazeCoT improves perceived explainability and trustworthiness by aligning MLLM social perception and social reasoning with human norms. We hope that GazeCoT, a versatile plug-and-play pipeline, can enable socially aware, MLLM-based HCI applications.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Omni-MMSI: Toward Identity-attributed Social Interaction UnderstandingXinpeng Li, Bolin Lai, Hardy Chen, Shijian Deng 等CVPR 2026 · 被引用 3 次
- Towards Aligning Multimodal LLMs with Human Experts: A Focus on Parent-Child InteractionWeiyan Shi, Kenny Tsu Wei ChooCHI 2026 · 被引用 2 次
- Video-Only ToM: Enhancing Theory of Mind in Multimodal Large Language ModelsSiqi Liu, Xinyang Li, Bochao Zou, Junbao Zhuo 等CVPR 2026
- Visual Intention Grounding for Egocentric AssistantsPengzhan Sun, Junbin Xiao, Tze Ho Elden Tse, Yicong Li 等ICCV 2025 · 被引用 2 次
- Seeing Conversations: Communication Context Identification in Egocentric VideoTobias Dorszewski, Jens HjortkjærCVPR 2026
