Can Vision-Language Models Answer Face to Face Questions in the Real-World?
Reza Pourreza, Rishit Dagli, Apratim Bhattacharyya, Sunny Panchal, Guillaume Berger, Roland Memisevic
Abstract
AI models have made significant strides in recent years in their ability to describe and answer questions about real-world images. They have also made progress in the ability to converse with users in real-time using audio input. This raises the question: have we reached the point where AI models, connected to a camera and microphone, can converse with users in real-time about scenes and events that are unfolding live in front of the camera? This has been a long-standing goal in AI and is a prerequisite for real-world AI assistants and humanoid robots to interact with humans in everyday situations. In this work, we introduce a new dataset and benchmark, the Qualcomm Interactive Video Dataset (IVD), which allows us to assess the extent to which existing models can support these abilities, and to what degree these capabilities can be instilled through fine-tuning. The dataset is based on a simple question-answering setup, where users ask questions that the system has to answer, in real-time, based on the camera and audio input. We show that existing models fall far behind human performance on this task, and we identify the main sources for the performance gap. However, we also show that for many of the required perceptual skills, fine-tuning on this form of data can significantly reduce this gap.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09cde461-253e-4f0d-80ac-a3a66e45597eCited by top-tier papers1
Ask how each one uses itBuilds on21
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
Related papers
- SIMMC-VR: A Task-oriented Multimodal Dialog Dataset with Situated and Immersive VR StreamsTe-Lin Wu, Satwik Kottur, Andrea Madotto, Mahmoud Azab et al.ACL 2023 · 4 citations
- GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented UnderstandingDongping Chen, Yue Huang, Siyuan Wu, Jingyu Tang et al.ICLR 2025 · 1 citation
- ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life VideosTe-Lin Wu, Zi-Yi Dou, Qingyuan Hu, Yu Hou et al.EMNLP 2023 · 3 citations
- VisAssist: A Visually Impaired-Captured Video Question Answering Benchmark for Assistive SystemsQi Gao, Heng Li, Yixin Zhou, Meixuan Zhou et al.AAAI 2026
- RIVER: A Real-Time Interaction Benchmark for Video LLMsYansong Shi, Qingsong Zhao, Tianxiang Jiang, Xiangyu Zeng et al.ICLR 2026 · 12 citations
