Visual Captions: Augmenting Verbal Communication with On-the-fly Visuals
Xingyu Bruce Liu, Vladimir Kirilyuk, Xiuxiu Yuan, Alex Olwal, Peggy Chi, Xiang 'Anthony' Chen, Ruofei Du
摘要
Video conferencing solutions like Zoom, Google Meet, and Microsoft Teams are becoming increasingly popular for facilitating conversations, and recent advancements such as live captioning help people better understand each other. We believe that the addition of visuals based on the context of conversations could further improve comprehension of complex or unfamiliar concepts. To explore the potential of such capabilities, we conducted a formative study through remote interviews (N=10) and crowdsourced a dataset of over 1500 sentence-visual pairs across a wide range of contexts. These insights informed Visual Captions, a real-time system that integrates with a video conferencing platform to enrich verbal communication. Visual Captions leverages a fine-tuned large language model to proactively suggest relevant visuals in open-vocabulary conversations. We present findings from a lab study (N=26) and an in-the-wild case study (N=10), demonstrating how Visual Captions can help improve communication through visual augmentation in various scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented RealityJaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu 等CHI 2024 · 被引用 86 次
- Proactive Conversational Agents with Inner ThoughtsXingyu Bruce Liu, Shitao Fang, Weiyan Shi, Chien-Sheng Wu 等CHI 2025 · 被引用 76 次
- Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature ReviewRock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas 等CHI 2025 · 被引用 51 次
- Memoro: Using Large Language Models to Realize a Concise Interface for Real-Time Memory AugmentationWazeer Deen Zulfikar, Samantha W. T. Chan, Pattie MaesCHI 2024 · 被引用 41 次
- ThingShare: Ad-Hoc Digital Copies of Physical Objects for Sharing Things in Video MeetingsErzhen Hu, Jens Emil Sloth Grønbæk, Wen Ying, Ruofei Du 等CHI 2023 · 被引用 30 次
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Improved Techniques for Training Score-Based Generative ModelsYang Song, Stefano ErmonNeurIPS 2020 · 被引用 1,527 次
- ILVR: Conditioning Method for Denoising Diffusion Probabilistic ModelsJooyoung Choi, Sungwon Kim, Yonghyun Jeong, Youngjune Gwon 等ICCV 2021 · 被引用 933 次
- SinDDM: A Single Image Denoising Diffusion ModelVladimir Kulikov, Shahar Yadin, Matan Kleiner, Tomer MichaeliICML 2023 · 被引用 113 次
相关 Paper
- Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio CaptioningSangyeon Cho, Mingi Kim, Jinkwon Hwang, Jaehoon Go 等EMNLP 2025
- CompCap: Improving Multimodal Large Language Models with Composite CaptionsXiaohui Chen, Satya Narayan Shukla, Mahmoud Azab, Aashu Singh 等ICCV 2025 · 被引用 2 次
- MM-SeR: Multimodal Self-Refinement for Lightweight Image CaptioningJunha Song, Yongsik Jo, So Yeon Min, Quanting Xie 等CVPR 2026
- Distilling Vision-Language Models on Millions of VideosYue Zhao, Long Zhao, Xingyi Zhou, Jialin Wu 等CVPR 2024
- VisText: A Benchmark for Semantically Rich Chart CaptioningBenny J. Tang, Angie W. Boggust, Arvind SatyanarayanACL 2023 · 被引用 44 次
