Generating Audio-Visual Slideshows from Text Articles Using Word Concreteness
Mackenzie Leake, Hijung Valentina Shin, Joy O. Kim, Maneesh Agrawala
摘要
We present a system that automatically transforms text articles into audio-visual slideshows by leveraging the notion of word concreteness, which measures how strongly a word or phrase is related to some perceptible concept. In a formative study we learn that people not only prefer such audio-visual slideshows but find that the content is easier to understand compared to text articles or text articles augmented with images. We use word concreteness to select search terms and find images relevant to the text. Then, based on the distribution of concrete words and the grammatical structure of an article, we time-align selected images with audio narration obtained through text-to-speech to produce audio-visual slideshows. In a user evaluation we find that our concreteness-based algorithm selects images that are highly relevant to the text. The quality of our slideshows is comparable to slideshows produced manually using standard video editing tools, and people strongly prefer our slideshows to those generated using a simple keyword-search based approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- MetaMap: Supporting Visual Metaphor Ideation through Multi-dimensional Example-based ExplorationYouwen Kang, Zhida Sun, Sitong Wang, Zeyu Huang 等CHI 2021 · 被引用 78 次
- GenAssist: Making Image Generation AccessibleMina Huh, Yi-Hao Peng, Amy PavelUIST 2023 · 被引用 58 次
- Visual Captions: Augmenting Verbal Communication with On-the-fly VisualsXingyu Bruce Liu, Vladimir Kirilyuk, Xiuxiu Yuan, Alex Olwal 等CHI 2023 · 被引用 45 次
- Crosspower: Bridging Graphics and LinguisticsHaijun XiaUIST 2020 · 被引用 32 次
- HelpViz: Automatic Generation of Contextual Visual Mobile Tutorials from Text-Based InstructionsMingyuan Zhong, Gang Li, Peggy Chi, Yang LiUIST 2021 · 被引用 27 次
相关 Paper
- Crosscast: Adding Visuals to Audio Travel PodcastsHaijun Xia, Jennifer Jacobs, Maneesh AgrawalaUIST 2020 · 被引用 44 次
- Automatic Instructional Video Creation from a Markdown-Formatted TutorialPeggy Chi, Nathan Frey, Katrina Panovich, Irfan EssaUIST 2021 · 被引用 27 次
- Word-As-Image for Semantic TypographyShir Iluz, Yael Vinker, Amir Hertz, Daniel Berio 等SIGGRAPH 2023 · 被引用 67 次
- RealityTalk: Real-Time Speech-Driven Augmented Presentation for AR Live StorytellingJian Liao, Adnan Karim, Shivesh Singh Jadon, Rubaiat Habib Kazi 等UIST 2022 · 被引用 44 次
- Unveiling the mystery of visual attributes of concrete and abstract concepts: Variability, nearest neighbors, and challenging categoriesTarun Tater, Sabine Schulte im Walde, Diego FrassinelliEMNLP 2024 · 被引用 2 次
