ChatCam: Embracing LLMs for Contextual Chatting-to-Camera with Interest-Oriented Video Summarization
Kaijie Xiao, Yi Gao, Fu Li, Weifeng Xu, Pengzhi Chen, Wei Dong
Abstract
Cameras are ubiquitous in society, with users increasingly looking to extract insights about the physical world. Current human-to-camera interaction methods, while advanced, still need to support an intuitive, conversational interaction as one would expect in human-to-human communication. To achieve a more natural interaction between humans and cameras, we proposed a novel contextual chatting-to-camera paradigm. This paradigm allows users to interact with the camera using natural language including raising interests and questions. In response, the camera can customize specific tasks tailored to these interests and attempt to provide answers to the questions asked. We designed ChatCam, embracing LLMs for contextual chatting-to-camera with interest-oriented video summarization. With a novel prompt with the actor-critic LLMs approach, ChatCam can understand users' interests and translate them into some tasks and objects. ChatCam can also customize relevant models with the help of the multi-modal large language model and deep reinforcement learning on the resource-constrained edge and maintain high accuracy. Results show that ChatCam achieves an improvement up to 43.9% in understanding user interests and 21.1% in model accuracy compared to state-of-the-art methods in multiple settings. Various examples and the user study also prove the effectiveness of ChatCam in practice.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get eb060106-0522-45c1-b9c1-5dafcbad1b33Cited by top-tier papers2
- Gestura: A LVLM-Powered System Bridging Motion and Semantics for Real-Time Free-Form Gesture UnderstandingZhuoming Li, Aitong Liu, Mengxi Jia, Yubo Lu et al.UbiComp 2026 · 1 citation
- The Pervasive Blind Spot: Benchmarking VLM Inference Risks on Everyday Personal VideosShuning Zhang, Zhaoxin Li, Changxi Wen, Ying Ma et al.UbiComp 2026
Related papers
- ChatIoT: Zero-code Generation of Trigger-action Based IoT ProgramsYi Gao, Kaijie Xiao, Fu Li, Weifeng Xu et al.UbiComp 2024 · 18 citations
- UniOVA: Universal On-demand Video Analytics with Edge-Cloud Collaborative Multimodal LLMKaijie Xiao, Yi Gao, Wei DongUbiComp 2026
- Agentic Video Summarization via Self-Reflecting Multimodal UnderstandingMiaotian Guo, Shuguang Dou, Yin Li, Aidong Men et al.CVPR 2026
- ChatCam: Empowering Camera Control through Conversational AIXinhang Liu, Yu-Wing Tai, Chi-Keung TangNeurIPS 2024 · 19 citations
- TikTalk: A Video-Based Dialogue Dataset for Multi-Modal Chitchat in Real WorldHongpeng Lin, Ludan Ruan, Wenke Xia, Peiyu Liu et al.ACM MM 2023 · 10 citations
