Crosscast: Adding Visuals to Audio Travel Podcasts
Haijun Xia, Jennifer Jacobs, Maneesh Agrawala
摘要
Audio travel podcasts are a valuable source of information for travelers. Yet, travel is, in many ways, a visual experience and the lack of visuals in travel podcasts can make it difficult for listeners to fully understand the places being discussed. We present Crosscast: a system for automatically adding visuals to audio travel podcasts. Given an audio travel podcast as input, Crosscast uses natural language processing and text mining to identify geographic locations and descriptive keywords within the podcast transcript. Crosscast then uses these locations and keywords to automatically select relevant photos from online repositories and synchronizes their display to align with the audio narration. In a user evaluation, we find that 85.7% of the participants preferred Crosscast generated audio-visual travel podcasts compared to audio-only travel podcasts.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper18
- Graphologue: Exploring Large Language Model Responses with Interactive DiagramsPeiling Jiang, Jude Rayan, Steven P. Dow, Haijun XiaUIST 2023 · 被引用 135 次
- PopBlends: Strategies for Conceptual Blending with Large Language ModelsSitong Wang, Savvas Petridis, Taeahn Kwon, Xiaojuan Ma 等CHI 2023 · 被引用 57 次
- DataParticles: Block-based and Language-oriented Authoring of Animated Unit VisualizationsYining Cao, Jane L. E, Chen Zhu-Tian, Haijun XiaCHI 2023 · 被引用 53 次
- Visual Captions: Augmenting Verbal Communication with On-the-fly VisualsXingyu Bruce Liu, Vladimir Kirilyuk, Xiuxiu Yuan, Alex Olwal 等CHI 2023 · 被引用 45 次
- Sporthesia: Augmenting Sports Videos Using Natural LanguageChen Zhu-Tian, Qisen Yang, Xiao Xie, Johanna Beyer 等IEEE VIS 2022 · 被引用 33 次
相关 Paper
- Generating Audio-Visual Slideshows from Text Articles Using Word ConcretenessMackenzie Leake, Hijung Valentina Shin, Joy O. Kim, Maneesh AgrawalaCHI 2020 · 被引用 36 次
- MoonCast: High-Quality Zero-Shot Podcast GenerationZeqian Ju, Dongchao Yang, Kai Shen, Yichong Leng 等NeurIPS 2025 · 被引用 32 次
- AutoAD II: The Sequel - Who, When, and What in Movie Audio DescriptionTengda Han, Max Bain, Arsha Nagrani, Gül Varol 等ICCV 2023 · 被引用 55 次
- VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual AugmentationsBaoquan Zhao, Xiaofan Ma, Qianshi Pang, Ruomei Wang 等ACM MM 2025 · 被引用 1 次
- Towards Abstractive Grounded Summarization of Podcast TranscriptsKaiqiang Song, Chen Li, Xiaoyang Wang, Dong Yu 等ACL 2022 · 被引用 11 次
