OmniScribe: Authoring Immersive Audio Descriptions for 360° Videos
Ruei-Che Chang, Chao-Hsien Ting, Chia-Sheng Hung, Wan-Chen Lee, Liang-Jin Chen, Yu-Tzu Chao, Bing-Yu Chen, Anhong Guo
Abstract
Blind people typically access videos via audio descriptions (AD) crafted by sighted describers who comprehend, select, and describe crucial visual content in the videos. 360° video is an emerging storytelling medium that enables immersive experiences that people may not possibly reach in everyday life. However, the omnidirectional nature of 360° videos makes it challenging for describers to perceive the holistic visual content and interpret spatial information that is essential to create immersive ADs for blind people. Through a formative study with a professional describer, we identified key challenges in describing 360° videos and iteratively designed OmniScribe, a system that supports the authoring of immersive ADs for 360° videos. OmniScribe uses AI-generated content-awareness overlays for describers to better grasp 360° video content. Furthermore, OmniScribe enables describers to author spatial AD and immersive labels for blind users to consume the videos immersively with our mobile prototype. In a study with 11 professional and novice describers, we demonstrated the value of OmniScribe in the authoring workflow; and a study with 8 blind participants revealed the promise of immersive AD over standard AD for 360° videos. Finally, we discuss the implications of promoting 360° video accessibility.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5edfeeeb-af3f-47d8-91d8-d50b9583e3a9Cited by top-tier papers11
- WorldScribe: Towards Context-Aware Live Visual DescriptionsRuei-Che Chang, Yuxuan Liu, Anhong GuoUIST 2024 · 54 citations
- Making Short-Form Videos Accessible with Hierarchical Video SummariesTess Van Daele, Akhil Iyer, Yuning Zhang, Jalyn C. Derry et al.CHI 2024 · 37 citations
- DanmuA11y: Making Time-Synced On-Screen Video Comments (Danmu) Accessible to Blind and Low Vision Users via Multi-Viewer Audio DiscussionsShuchang Xu, Xiaofu Jin, Huamin Qu, Yukang YanCHI 2025 · 26 citations
- "It's Kind of Context Dependent": Understanding Blind and Low Vision People's Video Accessibility Preferences Across Viewing ScenariosLucy Jiang, Crescentia Jung, Mahika Phutane, Abigale Stangl et al.CHI 2024 · 23 citations
- VideoA11y: Method and Dataset for Accessible Video DescriptionChaoyu Li, Sid Padmanabhuni, Maryam S. Cheema, Hasti Seifi et al.CHI 2025 · 23 citations
Builds on9
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Twitter A11y: A Browser Extension to Make Twitter Images AccessibleCole Gleason, Amy Pavel, Emma McCamey, Christina Low et al.CHI 2020 · 123 citations
- Virtual Reality Without Vision: A Haptic and Auditory White Cane to Navigate Complex Virtual WorldsAlexa F. Siu, Mike Sinclair, Robert Kovacs, Eyal Ofek et al.CHI 2020 · 115 citations
- What Makes Videos Accessible to Blind and Visually Impaired People?Xingyu Liu, Patrick Carrington, Xiang 'Anthony' Chen, Amy PavelCHI 2021 · 78 citations
- Rescribe: Authoring and Automatically Editing Audio DescriptionsAmy Pavel, Gabriel Reyes, Jeffrey P. BighamUIST 2020 · 72 citations
Related papers
- SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision ViewersZheng Ning, Brianna L. Wimer, Kaiwen Jiang, Keyi Chen et al.CHI 2024 · 27 citations
- Supporting Novices Author Audio Descriptions via Automatic FeedbackRosiana Natalie, Joshua Tseng, Hernisa Kacorri, Kotaro HaraCHI 2023 · 18 citations
- ADCanvas: Accessible and Conversational Audio Description Authoring for Blind and Low Vision CreatorsFranklin Mingzhe Li, Michael Xieyang Liu, Cynthia L. Bennett, Shaun K. KaneCHI 2026 · 2 citations
- Toward Automatic Audio Description Generation for Accessible VideosYujia Wang, Wei Liang, Haikun Huang, Yongqi Zhang et al.CHI 2021 · 86 citations
- CrossA11y: Identifying Video Accessibility Issues via Cross-modal GroundingXingyu Bruce Liu, Ruolin Wang, Dingzeyu Li, Xiang Anthony Chen et al.UIST 2022 · 24 citations
