GazeCoT: Unleashing Social Intelligence in Multimodal LLMs With Gaze-Informed Chain-of-Thought Reasoning
Zhoutong Ye, Xutong Wang, Chengwen Zhang, Ruiwen Zhang, Mingze Sun, Qinwei Li, Chun Yu, Yuanchun Shi
Abstract
Social intelligence is vital for effective human-AI interaction. While LLMs demonstrate strong text-based social intelligence, the vision modality remains challenging due to the presence of non-verbal social cues. For example, gaze is the primary conveyor of social attention, yet it cannot be accurately perceived and understood by multimodal LLMs (MLLMs). Therefore, we propose GazeCoT, a pipeline using gaze estimation models to provide MLLMs with the attention of people in images or videos. The gaze information is provided as visual and text prompts compiled into a structured context to support MLLM social reasoning. Benchmark evaluation confirms that GazeCoT enhances MLLMs’ social intelligence by improving gaze perception. A user study in a challenging application involving parent-child interactions demonstrates that GazeCoT improves perceived explainability and trustworthiness by aligning MLLM social perception and social reasoning with human norms. We hope that GazeCoT, a versatile plug-and-play pipeline, can enable socially aware, MLLM-based HCI applications.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b7d0e3c4-551a-4f5f-bb53-e111797bb7eeRelated papers
- Omni-MMSI: Toward Identity-attributed Social Interaction UnderstandingXinpeng Li, Bolin Lai, Hardy Chen, Shijian Deng et al.CVPR 2026 · 3 citations
- Towards Aligning Multimodal LLMs with Human Experts: A Focus on Parent-Child InteractionWeiyan Shi, Kenny Tsu Wei ChooCHI 2026 · 2 citations
- Video-Only ToM: Enhancing Theory of Mind in Multimodal Large Language ModelsSiqi Liu, Xinyang Li, Bochao Zou, Junbao Zhuo et al.CVPR 2026
- Visual Intention Grounding for Egocentric AssistantsPengzhan Sun, Junbin Xiao, Tze Ho Elden Tse, Yicong Li et al.ICCV 2025 · 2 citations
- Seeing Conversations: Communication Context Identification in Egocentric VideoTobias Dorszewski, Jens HjortkjærCVPR 2026
