SonifyAR: Context-Aware Sound Generation in Augmented Reality
Xia Su, Jon E. Froehlich, Eunyee Koh, Chang Xiao
Abstract
Sound plays a crucial role in enhancing user experience and immersiveness in Augmented Reality (AR). However, current platforms lack support for AR sound authoring due to limited interaction types, challenges in collecting and specifying context information, and difficulty in acquiring matching sound assets. We present SonifyAR, an LLM-based AR sound authoring system that generates context-aware sound effects for AR experiences. SonifyAR expands the current design space of AR sound and implements a Programming by Demonstration (PbD) pipeline to automatically collect contextual information of AR events, including virtual-content-semantics and real-world context. This context information is then processed by a large language model to acquire sound effects with Recommendation, Retrieval, Generation, and Transfer methods. To evaluate the usability and performance of our system, we conducted a user study with eight participants and created five example applications, including an AR-based science experiment, and an assistive application for low-vision AR users.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 10e092ac-857c-4608-a463-e0847b9b534eCited by top-tier papers9
- ImaginateAR: AI-Assisted In-Situ Authoring in Augmented RealityJaewook Lee, Filippo Aleotti, Diego Mazala, Guillermo Garcia-Hernando et al.UIST 2025 · 15 citations
- StreetViewAI: Making Street View Accessible Using Context-Aware Multimodal AIJon E. Froehlich, Alexander J. Fiannaca, Nimer Jaber, Victor Tsaran et al.UIST 2025 · 5 citations
- Scene2Hap: Generating Scene-Wide Haptics for VR from Scene Context with Multimodal LLMsArata Jingu, Easa AliAbbasi, Sara Safaee, Paul Strohmeier et al.CHI 2026 · 4 citations
- GestureCoach: Rehearsing for Engaging Talks with LLM-Driven Gesture RecommendationsAshwin Ram, Varsha Suresh, Artin Saberpour Abadian, Vera Demberg et al.UIST 2025 · 2 citations
- PAVAS: Physics-Aware Video-to-Audio SynthesisOh Hyun-Bin, Yuhta Takida, Toshimitsu Uesaka, Tae-Hyun Oh et al.CVPR 2026 · 2 citations
Builds on14
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao et al.ICLR 2021 · 1,902 citations
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li et al.NeurIPS 2023 · 1,778 citations
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei et al.ICML 2023 · 773 citations
- VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and DatasetSihan Chen, Handong Li, Qunbo Wang, Zijia Zhao et al.NeurIPS 2023 · 246 citations
- RealitySketch: Embedding Responsive Graphics and Visualizations in AR through Dynamic SketchingRyo Suzuki, Rubaiat Habib Kazi, Li-Yi Wei, Stephen DiVerdi et al.UIST 2020 · 98 citations
Related papers
- Exploring Large Language Model-Driven Agents for Environment-Aware Spatial Interactions and Conversations in Virtual Reality Role-Play ScenariosZiming Li, Huadong Zhang, Chao Peng, Roshan L. PeirisIEEE VR 2025 · 18 citations
- SocialMind: LLM-based Proactive AR Social Assistive System with Human-like Perception for In-situ Live InteractionsBufang Yang, Yunqi Guo, Lilin Xu, Zhenyu Yan et al.UbiComp 2025 · 26 citations
- GenAssist: Interactive Prompt-Driven XR Program GenerationSruti Srinidhi, Akul Singh, Edward Lu, Anthony RoweIEEE VR 2026
- ARify: Leveraging Narrated Instructional Videos to Create Augmented Reality Tutorials for Procedural TasksXiyun Hu, Chenfei Zhu, Shao-Kang Hsia, Dizhi Ma et al.CHI 2026 · 1 citation
- Satori 悟り: Towards Proactive AR Assistant with Belief-Desire-Intention User ModelingChenyi Li, Guande Wu, Gromit Yeuk-Yin Chan, Dishita G. Turakhia et al.CHI 2025 · 49 citations
