VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
Baoquan Zhao, Xiaofan Ma, Qianshi Pang, Ruomei Wang, Fan Zhou, Shujin Lin
Abstract
The widespread adoption of digital technology has ushered in a new era of digital transformation across all aspects of our lives. Online learning, social, and work activities, such as distance education, videoconferencing, interviews, and talks, have led to a dramatic increase in speech-rich video content. In contrast to other video types, such as surveillance footage, which typically contain abundant visual cues, speech-rich videos convey most of their meaningful information through the audio channel. This poses challenges for improving content consumption using existing visual-based video summarization, navigation, and exploration systems. In this paper, we present VisAug, a novel interactive system designed to enhance speech-rich video navigation and engagement by automatically generating informative and expressive visual augmentations based on the speech content of videos. Our findings suggest that this system has the potential to significantly enhance the consumption and engagement of information in an increasingly video-driven digital landscape.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b0c9470-14a9-45bb-95d3-9cdcba8cb12fBuilds on19
- CogView2: Faster and Better Text-to-Image Generation via Hierarchical TransformersMing Ding, Wendi Zheng, Wenyi Hong, Jie TangNeurIPS 2022 · 432 citations
- Differentiable Prompt Makes Pre-trained Language Models Better Few-shot LearnersNingyu Zhang, Luoqiu Li, Xiang Chen, Shumin Deng et al.ICLR 2022 · 205 citations
- An Information-theoretic Approach to Prompt Engineering Without Ground Truth LabelsTaylor Sorensen, Joshua Robinson, Christopher Michael Rytting, Alexander Glenn Shaw et al.ACL 2022 · 142 citations
- Expressive Text-to-Image Generation with Rich TextSongwei Ge, Taesung Park, Jun-Yan Zhu, Jia-Bin HuangICCV 2023 · 102 citations
- F2GAN: Fusing-and-Filling GAN for Few-shot Image GenerationYan Hong, Li Niu, Jianfu Zhang, Weijie Zhao et al.ACM MM 2020 · 93 citations
Related papers
- RealityTalk: Real-Time Speech-Driven Augmented Presentation for AR Live StorytellingJian Liao, Adnan Karim, Shivesh Singh Jadon, Rubaiat Habib Kazi et al.UIST 2022 · 44 citations
- Sporthesia: Augmenting Sports Videos Using Natural LanguageChen Zhu-Tian, Qisen Yang, Xiao Xie, Johanna Beyer et al.IEEE VIS 2022 · 33 citations
- Papeos: Augmenting Research Papers with Talk VideosTae Soo Kim, Matt Latzke, Jonathan Bragg, Amy X. Zhang et al.UIST 2023 · 17 citations
- SpeechMirror: A Multimodal Visual Analytics System for Personalized Reflection of Online Public Speaking EffectivenessZe-Yuan Huang, Qiang He, Kevin T. Maher, Xiaoming Deng et al.IEEE VIS 2023 · 7 citations
- SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision ViewersZheng Ning, Brianna L. Wimer, Kaiwen Jiang, Keyi Chen et al.CHI 2024 · 27 citations
