Visual Captions: Augmenting Verbal Communication with On-the-fly Visuals
Xingyu Bruce Liu, Vladimir Kirilyuk, Xiuxiu Yuan, Alex Olwal, Peggy Chi, Xiang 'Anthony' Chen, Ruofei Du
Abstract
Video conferencing solutions like Zoom, Google Meet, and Microsoft Teams are becoming increasingly popular for facilitating conversations, and recent advancements such as live captioning help people better understand each other. We believe that the addition of visuals based on the context of conversations could further improve comprehension of complex or unfamiliar concepts. To explore the potential of such capabilities, we conducted a formative study through remote interviews (N=10) and crowdsourced a dataset of over 1500 sentence-visual pairs across a wide range of contexts. These insights informed Visual Captions, a real-time system that integrates with a video conferencing platform to enrich verbal communication. Visual Captions leverages a fine-tuned large language model to proactively suggest relevant visuals in open-vocabulary conversations. We present findings from a lab study (N=26) and an in-the-wild case study (N=10), demonstrating how Visual Captions can help improve communication through visual augmentation in various scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d78990e-eea0-4709-ba3d-a508a973b630Cited by top-tier papers25
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented RealityJaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu et al.CHI 2024 · 86 citations
- Proactive Conversational Agents with Inner ThoughtsXingyu Bruce Liu, Shitao Fang, Weiyan Shi, Chien-Sheng Wu et al.CHI 2025 · 76 citations
- Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature ReviewRock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas et al.CHI 2025 · 51 citations
- Memoro: Using Large Language Models to Realize a Concise Interface for Real-Time Memory AugmentationWazeer Deen Zulfikar, Samantha W. T. Chan, Pattie MaesCHI 2024 · 41 citations
- ThingShare: Ad-Hoc Digital Copies of Physical Objects for Sharing Things in Video MeetingsErzhen Hu, Jens Emil Sloth Grønbæk, Wen Ying, Ruofei Du et al.CHI 2023 · 30 citations
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Improved Techniques for Training Score-Based Generative ModelsYang Song, Stefano ErmonNeurIPS 2020 · 1,527 citations
- ILVR: Conditioning Method for Denoising Diffusion Probabilistic ModelsJooyoung Choi, Sungwon Kim, Yonghyun Jeong, Youngjune Gwon et al.ICCV 2021 · 933 citations
- SinDDM: A Single Image Denoising Diffusion ModelVladimir Kulikov, Shahar Yadin, Matan Kleiner, Tomer MichaeliICML 2023 · 113 citations
Related papers
- Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio CaptioningSangyeon Cho, Mingi Kim, Jinkwon Hwang, Jaehoon Go et al.EMNLP 2025
- CompCap: Improving Multimodal Large Language Models with Composite CaptionsXiaohui Chen, Satya Narayan Shukla, Mahmoud Azab, Aashu Singh et al.ICCV 2025 · 2 citations
- MM-SeR: Multimodal Self-Refinement for Lightweight Image CaptioningJunha Song, Yongsik Jo, So Yeon Min, Quanting Xie et al.CVPR 2026
- Distilling Vision-Language Models on Millions of VideosYue Zhao, Long Zhao, Xingyi Zhou, Jialin Wu et al.CVPR 2024
- VisText: A Benchmark for Semantically Rich Chart CaptioningBenny J. Tang, Angie W. Boggust, Arvind SatyanarayanACL 2023 · 44 citations
