Crosscast: Adding Visuals to Audio Travel Podcasts
Haijun Xia, Jennifer Jacobs, Maneesh Agrawala
Abstract
Audio travel podcasts are a valuable source of information for travelers. Yet, travel is, in many ways, a visual experience and the lack of visuals in travel podcasts can make it difficult for listeners to fully understand the places being discussed. We present Crosscast: a system for automatically adding visuals to audio travel podcasts. Given an audio travel podcast as input, Crosscast uses natural language processing and text mining to identify geographic locations and descriptive keywords within the podcast transcript. Crosscast then uses these locations and keywords to automatically select relevant photos from online repositories and synchronizes their display to align with the audio narration. In a user evaluation, we find that 85.7% of the participants preferred Crosscast generated audio-visual travel podcasts compared to audio-only travel podcasts.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers18
- Graphologue: Exploring Large Language Model Responses with Interactive DiagramsPeiling Jiang, Jude Rayan, Steven P. Dow, Haijun XiaUIST 2023 · 135 citations
- PopBlends: Strategies for Conceptual Blending with Large Language ModelsSitong Wang, Savvas Petridis, Taeahn Kwon, Xiaojuan Ma et al.CHI 2023 · 57 citations
- DataParticles: Block-based and Language-oriented Authoring of Animated Unit VisualizationsYining Cao, Jane L. E, Chen Zhu-Tian, Haijun XiaCHI 2023 · 53 citations
- Visual Captions: Augmenting Verbal Communication with On-the-fly VisualsXingyu Bruce Liu, Vladimir Kirilyuk, Xiuxiu Yuan, Alex Olwal et al.CHI 2023 · 45 citations
- Sporthesia: Augmenting Sports Videos Using Natural LanguageChen Zhu-Tian, Qisen Yang, Xiao Xie, Johanna Beyer et al.IEEE VIS 2022 · 33 citations
Related papers
- Generating Audio-Visual Slideshows from Text Articles Using Word ConcretenessMackenzie Leake, Hijung Valentina Shin, Joy O. Kim, Maneesh AgrawalaCHI 2020 · 36 citations
- MoonCast: High-Quality Zero-Shot Podcast GenerationZeqian Ju, Dongchao Yang, Kai Shen, Yichong Leng et al.NeurIPS 2025 · 32 citations
- AutoAD II: The Sequel - Who, When, and What in Movie Audio DescriptionTengda Han, Max Bain, Arsha Nagrani, Gül Varol et al.ICCV 2023 · 55 citations
- VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual AugmentationsBaoquan Zhao, Xiaofan Ma, Qianshi Pang, Ruomei Wang et al.ACM MM 2025 · 1 citation
- Towards Abstractive Grounded Summarization of Podcast TranscriptsKaiqiang Song, Chen Li, Xiaoyang Wang, Dong Yu et al.ACL 2022 · 11 citations
