SoundCount: Sound Counting from Raw Audio with Dyadic Decomposition Neural Network
Yuhang He, Zhuangzhuang Dai, Niki Trigoni, Long Chen, Andrew Markham
摘要
In this paper, we study an underexplored, yet important and challenging problem: counting the number of distinct sounds in raw audio characterized by a high degree of polyphonicity. We do so by systematically proposing a novel end-to-end trainable neural network (which we call DyDecNet, consisting of a dyadic decomposition front-end and backbone network), and quantifying the difficulty level of counting depending on sound polyphonicity. The dyadic decomposition front-end progressively decomposes the raw waveform dyadically along the frequency axis to obtain time-frequency representation in multi-stage, coarse-to-fine manner. Each intermediate waveform convolved by a parent filter is further processed by a pair of child filters that evenly split the parent filter's carried frequency response, with the higher-half child filter encoding the detail and lower-half child filter encoding the approximation. We further introduce an energy gain normalization to normalize sound loudness variance and spectrum overlap, and apply it to each intermediate parent waveform before feeding it to the two child filters. To better quantify sound counting difficulty level, we further design three polyphony-aware metrics: polyphony ratio, max polyphony and mean polyphony. We test DyDecNet on various datasets to show its superiority, and we further show dyadic decomposition network can be used as a general front-end to tackle other acoustic tasks. Code: github.com/ yuhanghe01/SoundCount.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper1
相关 Paper
- Repetitive Activity Counting by Sight and SoundYunhua Zhang, Ling Shao, Cees G. M. SnoekCVPR 2021
- Polyphony: Diffusion-based Dual-Hand Action Segmentation with Alternating Vision Transformer and Semantic ConditioningHao Zheng, Hu Wang, Tiantian Zheng, Prajjwal Bhattarai 等CVPR 2026 · 被引用 2 次
- Recursive Visual Sound Separation Using Minus-Plus NetXudong Xu, Bo Dai, Dahua LinICCV 2019 · 被引用 95 次
- FlowDec: A flow-based full-band general audio codec with high perceptual qualitySimon Welker, Matthew Le, Ricky T. Q. Chen, Wei-Ning Hsu 等ICLR 2025
- Time-Frequency Domain Fusion Enhancement for Audio Super-ResolutionYe Tian, Zhe Wang, Jianguo Sun, Liguo ZhangACM MM 2024 · 被引用 1 次
