SoundDet: Polyphonic Moving Sound Event Detection and Localization from Raw Waveform
Yuhang He, Niki Trigoni, Andrew Markham
摘要
We present a new framework SoundDet, which is an end-to-end trainable and light-weight framework, for polyphonic moving sound event detection and localization. Prior methods typically approach this problem by preprocessing raw waveform into time-frequency representations, which is more amenable to process with wellestablished image processing pipelines. Prior methods also detect in segment-wise manner, leading to incomplete and partial detections. Sound-Det takes a novel approach and directly consumes the raw, multichannel waveform and treats the spatio-temporal sound event as a complete "sound-object" to be detected. Specifically, Sound-Det consists of a backbone neural network and two parallel heads for temporal detection and spatial localization, respectively. Given the large sampling rate of raw waveform, the backbone network first learns a set of phase-sensitive and frequency-selective bank of filters to explicitly retain direction-of-arrival information, whilst being highly computationally and parametrically efficient than standard 1D/2D convolution. A dense sound event proposal map is then constructed to handle the challenges of predicting events with large varying temporal duration. Accompanying the dense proposal map are a temporal overlapness map and a motion smoothness map that measure a proposal's confidence to be an event from temporal detection accuracy and movement consistency perspective. Involving the two maps guarantees SoundDet to be trained in a spatiotemporally unified manner. Experimental results on the public DCASE dataset show the advantage of SoundDet on both segment-based and our newly proposed event-based evaluation system.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Deep Neural Room Acoustics PrimitiveYuhang He, Anoop Cherian, Gordon Wichern, Andrew MarkhamICML 2024 · 被引用 6 次
- SoundCount: Sound Counting from Raw Audio with Dyadic Decomposition Neural NetworkYuhang He, Zhuangzhuang Dai, Niki Trigoni, Long Chen 等AAAI 2024 · 被引用 3 次
- DeepASA: An Object-Oriented Multi-Purpose Network for Auditory Scene AnalysisDongheon Lee, Younghoo Kwon, Jung-Woo ChoiNeurIPS 2025 · 被引用 3 次
- RiTTA: Modeling Event Relations in Text-to-Audio GenerationYuhang He, Yash Jain, Xubo Liu, Andrew Markham 等EMNLP 2025 · 被引用 1 次
- Aurelius: Relation Aware Text-to-Audio Generation At ScaleYuhang He, He Liang, Yash Jain, Andrew Markham 等ICLR 2026
相关 Paper
- Enhanced Audio Tagging via Multi- to Single-Modal Teacher-Student Mutual LearningYifang Yin, Harsh Shrivastava, Ying Zhang, Zhenguang Liu 等AAAI 2021 · 被引用 18 次
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 被引用 271 次
- Adaptive Hierarchical Pooling for Weakly-supervised Sound Event DetectionLijian Gao, Ling Zhou, Qirong Mao, Ming DongACM MM 2022 · 被引用 6 次
- UniCon: Unified Context Network for Robust Active Speaker DetectionYuanhang Zhang, Susan Liang, Shuang Yang, Xiao Liu 等ACM MM 2021 · 被引用 40 次
- Span-based Audio-Visual LocalizationYiling Wu, Xinfeng Zhang, Yaowei Wang, Qingming HuangACM MM 2022 · 被引用 7 次
