Simulating Human Audiovisual Search Behavior
Hyunsung Cho, Xuejing Luo, Byungjoo Lee, David Lindlbauer, Antti Oulasvirta
Abstract
Locating a target based on auditory and visual cues—such as finding a car in a crowded parking lot or identifying a speaker in a virtual meeting—requires balancing effort, time, and accuracy under uncertainty. Existing models of audiovisual search often treat perception and action in isolation, overlooking how people adaptively coordinate movement and sensory strategies. We present Sensonaut, a computational model of embodied audiovisual search. The core assumption is that people deploy their body and sensory systems in ways they believe will most efficiently improve their chances of locating a target, trading off time and effort under perceptual constraints. Our model formulates this as a resource-rational decision-making problem under partial observability. We validate the model against newly collected human data, showing that it reproduces both adaptive scaling of search time and effort under task complexity, occlusion, and distraction, and characteristic human errors. Our simulation of human-like resource-rational search informs the design of audiovisual interfaces that minimize search cost and cognitive load.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ab01449-c141-4a28-aae9-df6be0ba46a3Builds on15
- AUIT - the Adaptive User Interfaces Toolkit for Designing XR ApplicationsJoão Marcelo Evangelista Belo, Mathias N. Lystbæk, Anna Maria Feit, Ken Pfeuffer et al.UIST 2022 · 114 citations
- An Adaptive Model of Gaze-based SelectionXiuli Chen, Aditya Acharya, Antti Oulasvirta, Andrew HowesCHI 2021 · 38 citations
- CRTypist: Simulating Touchscreen Typing Behavior via Computational RationalityDanqing Shi, Yujun Zhu, Jussi P. P. Jokinen, Aditya Acharya et al.CHI 2024 · 26 citations
- Modeling Human Exploration Through Resource-Rational Reinforcement LearningMarcel Binz, Eric SchulzNeurIPS 2022 · 23 citations
- Sensorimotor Simulation of Redirected Reaching using Stochastic Optimal Feedback ControlEric J. Gonzalez, Sean FollmerCHI 2023 · 21 citations
Related papers
- CAVEN: An Embodied Conversational Agent for Efficient Audio-Visual Navigation in Noisy EnvironmentsXiulong Liu, Sudipta Paul, Moitreya Chatterjee, Anoop CherianAAAI 2024 · 16 citations
- SeekUI: Predicting Visual Search Behavior on Graphical User Interfaces with a Reward-Augmented Vision Language ModelZixin Guo, Yue Jiang, Luis A. Leiva, Antti OulasvirtaCHI 2026 · 1 citation
- Embodied Amodal Recognition: Learning to Move to Perceive ObjectsJianwei Yang, Zhile Ren, Mingze Xu, Xinlei Chen et al.ICCV 2019 · 70 citations
- NaVLA: A Vision-Language-Audio-Action Model for Multimodal Instruction NavigationJugang Fan, Peihao Chen, Changhao Li, Qing Du et al.AAAI 2026
- AVLEN: Audio-Visual-Language Embodied Navigation in 3D EnvironmentsSudipta Paul, Amit Roy-Chowdhury, Anoop CherianNeurIPS 2022 · 43 citations
