Generating Audio-Visual Slideshows from Text Articles Using Word Concreteness
Mackenzie Leake, Hijung Valentina Shin, Joy O. Kim, Maneesh Agrawala
Abstract
We present a system that automatically transforms text articles into audio-visual slideshows by leveraging the notion of word concreteness, which measures how strongly a word or phrase is related to some perceptible concept. In a formative study we learn that people not only prefer such audio-visual slideshows but find that the content is easier to understand compared to text articles or text articles augmented with images. We use word concreteness to select search terms and find images relevant to the text. Then, based on the distribution of concrete words and the grammatical structure of an article, we time-align selected images with audio narration obtained through text-to-speech to produce audio-visual slideshows. In a user evaluation we find that our concreteness-based algorithm selects images that are highly relevant to the text. The quality of our slideshows is comparable to slideshows produced manually using standard video editing tools, and people strongly prefer our slideshows to those generated using a simple keyword-search based approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0bbdd45b-651c-486c-a3c0-99338792712eCited by top-tier papers11
- MetaMap: Supporting Visual Metaphor Ideation through Multi-dimensional Example-based ExplorationYouwen Kang, Zhida Sun, Sitong Wang, Zeyu Huang et al.CHI 2021 · 78 citations
- GenAssist: Making Image Generation AccessibleMina Huh, Yi-Hao Peng, Amy PavelUIST 2023 · 58 citations
- Visual Captions: Augmenting Verbal Communication with On-the-fly VisualsXingyu Bruce Liu, Vladimir Kirilyuk, Xiuxiu Yuan, Alex Olwal et al.CHI 2023 · 45 citations
- Crosspower: Bridging Graphics and LinguisticsHaijun XiaUIST 2020 · 32 citations
- HelpViz: Automatic Generation of Contextual Visual Mobile Tutorials from Text-Based InstructionsMingyuan Zhong, Gang Li, Peggy Chi, Yang LiUIST 2021 · 27 citations
Related papers
- Crosscast: Adding Visuals to Audio Travel PodcastsHaijun Xia, Jennifer Jacobs, Maneesh AgrawalaUIST 2020 · 44 citations
- Automatic Instructional Video Creation from a Markdown-Formatted TutorialPeggy Chi, Nathan Frey, Katrina Panovich, Irfan EssaUIST 2021 · 27 citations
- Word-As-Image for Semantic TypographyShir Iluz, Yael Vinker, Amir Hertz, Daniel Berio et al.SIGGRAPH 2023 · 67 citations
- RealityTalk: Real-Time Speech-Driven Augmented Presentation for AR Live StorytellingJian Liao, Adnan Karim, Shivesh Singh Jadon, Rubaiat Habib Kazi et al.UIST 2022 · 44 citations
- Unveiling the mystery of visual attributes of concrete and abstract concepts: Variability, nearest neighbors, and challenging categoriesTarun Tater, Sabine Schulte im Walde, Diego FrassinelliEMNLP 2024 · 2 citations
