VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
Baoquan Zhao, Xiaofan Ma, Qianshi Pang, Ruomei Wang, Fan Zhou, Shujin Lin
摘要
The widespread adoption of digital technology has ushered in a new era of digital transformation across all aspects of our lives. Online learning, social, and work activities, such as distance education, videoconferencing, interviews, and talks, have led to a dramatic increase in speech-rich video content. In contrast to other video types, such as surveillance footage, which typically contain abundant visual cues, speech-rich videos convey most of their meaningful information through the audio channel. This poses challenges for improving content consumption using existing visual-based video summarization, navigation, and exploration systems. In this paper, we present VisAug, a novel interactive system designed to enhance speech-rich video navigation and engagement by automatically generating informative and expressive visual augmentations based on the speech content of videos. Our findings suggest that this system has the potential to significantly enhance the consumption and engagement of information in an increasingly video-driven digital landscape.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- CogView2: Faster and Better Text-to-Image Generation via Hierarchical TransformersMing Ding, Wendi Zheng, Wenyi Hong, Jie TangNeurIPS 2022 · 被引用 432 次
- Differentiable Prompt Makes Pre-trained Language Models Better Few-shot LearnersNingyu Zhang, Luoqiu Li, Xiang Chen, Shumin Deng 等ICLR 2022 · 被引用 205 次
- An Information-theoretic Approach to Prompt Engineering Without Ground Truth LabelsTaylor Sorensen, Joshua Robinson, Christopher Michael Rytting, Alexander Glenn Shaw 等ACL 2022 · 被引用 142 次
- Expressive Text-to-Image Generation with Rich TextSongwei Ge, Taesung Park, Jun-Yan Zhu, Jia-Bin HuangICCV 2023 · 被引用 102 次
- F2GAN: Fusing-and-Filling GAN for Few-shot Image GenerationYan Hong, Li Niu, Jianfu Zhang, Weijie Zhao 等ACM MM 2020 · 被引用 93 次
相关 Paper
- RealityTalk: Real-Time Speech-Driven Augmented Presentation for AR Live StorytellingJian Liao, Adnan Karim, Shivesh Singh Jadon, Rubaiat Habib Kazi 等UIST 2022 · 被引用 44 次
- Sporthesia: Augmenting Sports Videos Using Natural LanguageChen Zhu-Tian, Qisen Yang, Xiao Xie, Johanna Beyer 等IEEE VIS 2022 · 被引用 33 次
- Papeos: Augmenting Research Papers with Talk VideosTae Soo Kim, Matt Latzke, Jonathan Bragg, Amy X. Zhang 等UIST 2023 · 被引用 17 次
- SpeechMirror: A Multimodal Visual Analytics System for Personalized Reflection of Online Public Speaking EffectivenessZe-Yuan Huang, Qiang He, Kevin T. Maher, Xiaoming Deng 等IEEE VIS 2023 · 被引用 7 次
- SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision ViewersZheng Ning, Brianna L. Wimer, Kaiwen Jiang, Keyi Chen 等CHI 2024 · 被引用 27 次
