BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural Speech Synthesis with Flow Matching Models
Susan Liang, Dejan Markovic, Israel D. Gebru, Steven Krenn, Todd Keebler, Jacob Sandakly, Frank Yu, Samuel Hassel, Chenliang Xu, Alexander Richard
Abstract
Binaural rendering aims to synthesize binaural audio that mimics natural hearing based on a mono audio and the locations of the speaker and listener. Although many methods have been proposed to solve this problem, they struggle with rendering quality and streamable inference. Synthesizing high-quality binaural audio that is indistinguishable from real-world recordings requires precise modeling of binaural cues, room reverb, and ambient sounds. Additionally, real-world applications demand streaming inference. To address these challenges, we propose a flow matching based streaming binaural speech synthesis framework called BinauralFlow. We consider binaural rendering to be a generation problem rather than a regression problem and design a conditional flow matching model to render high-quality audio. Moreover, we design a causal U-Net architecture that estimates the current audio frame solely based on past information to tailor generative models for streaming inference. Finally, we introduce a continuous inference pipeline incorporating streaming STFT/ISTFT operations, a buffer bank, a midpoint solver, and an early skip schedule to improve rendering continuity and speed. Quantitative and qualitative evaluations demonstrate the superiority of our method over SOTA approaches. A perceptual study further reveals that our model is nearly indistinguishable from realworld recordings, with a 42% confusion rate. We recommend that readers visit our project page for demo videos: https://liangsusan-git. github.io/project/binauralflow/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 06c3a2b4-8009-43ea-b1f4-9f0a92d62ffdCited by top-tier papers4
- SmartDJ: Declarative Audio Editing with Audio Language ModelZitong Lan, Yiduo Hao, Mingmin ZhaoICLR 2026 · 11 citations
- -AVAS: Can Physics-Integrated Audio-Visual Modeling Boost Neural Acoustic Synthesis?Susan Liang, Chao Huang, Yunlong Tang, Zeliang Zhang et al.ICCV 2025 · 4 citations
- Few-shot Acoustic Synthesis with Multimodal Flow MatchingAmandine BrunettoCVPR 2026 · 2 citations
- Physics-Informed Audio-Geometry-Grid Representation Learning for Universal Sound Source LocalizationMin-Sang Baek, Gyeong-Su Kim, Donghyun Kim, Joon-Hyuk ChangICLR 2026
Builds on17
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao et al.ICLR 2021 · 1,902 citations
Related papers
- Neural Synthesis of Binaural Speech From Mono AudioAlexander Richard, Dejan Markovic, Israel D. Gebru, Steven Krenn et al.ICLR 2021 · 73 citations
- Listen2Scene: Interactive material-aware binaural sound propagation for reconstructed 3D scenesAnton Ratnarajah, Dinesh ManochaIEEE VR 2024 · 11 citations
- BinauralGrad: A Two-Stage Conditional Diffusion Probabilistic Model for Binaural Audio SynthesisYichong Leng, Zehua Chen, Junliang Guo, Haohe Liu et al.NeurIPS 2022 · 86 citations
- Localize to Binauralize: Audio Spatialization from Visual Sound Source LocalizationKranthi Kumar Rachavarapu, Aakanksha, Vignesh Sundaresha, A. N. RajagopalanICCV 2021 · 27 citations
- Cinematic Audio Source Separation Using Visual CuesKang Zhang, Suyeon Lee, Arda Senocak, Joon Son ChungCVPR 2026 · 1 citation
