Deep Audio Priors Emerge From Harmonic Convolutional Networks
Zhoutong Zhang, Yunyun Wang, Chuang Gan, Jiajun Wu, Joshua B. Tenenbaum, Antonio Torralba, William T. Freeman
Abstract
Convolutional neural networks (CNNs) excel in image recognition and generation. Among many efforts to explain their effectiveness, experiments show that CNNs carry strong inductive biases that capture natural image priors. Do deep networks also have inductive biases for audio signals? In this paper, we empirically show that current network architectures for audio processing do not show strong evidence in capturing such priors. We propose Harmonic Convolution, an operation that helps deep networks distill priors in audio signals by explicitly utilizing the harmonic structure within. This is done by engineering the kernel to be supported by sets of harmonic series, instead of local neighborhoods for convolutional kernels. We show that networks using Harmonic Convolution can reliably model audio priors and achieve high performance in unsupervised audio restoration tasks. With Harmonic Convolution, they also achieve better generalization performance for sound source separation.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6c17204b-efcd-4491-bb73-4edcf7c4795dCited by top-tier papers3
- Listening to Sounds of Silence for Speech DenoisingRuilin Xu, Rundi Wu, Yuko Ishiwaka, Carl Vondrick et al.NeurIPS 2020 · 40 citations
- Catch-A-Waveform: Learning to Generate Audio from a Single Short ExampleGal Greshler, Tamar Rott Shaham, Tomer MichaeliNeurIPS 2021 · 29 citations
- Deep Harmonic Finesse: Signal Separation in Wearable Systems with Limited DataMahya Saffarpour, Weitai Qian, Kourosh Vali, Begum Kasap et al.DAC 2024 · 4 citations
Related papers
- UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image RecognitionXiaohan Ding, Yiyuan Zhang, Yixiao Ge, Sijie Zhao et al.CVPR 2024
- SinBasis Networks: Matrix-Equivalent Feature Extraction for Wave-Like Optical SpectrogramsYuzhou Zhu, Zheng Zhang, Ruyi Zhang, Liang ZhouAAAI 2026 · 1 citation
- LEAF: A Learnable Frontend for Audio ClassificationNeil Zeghidour, Olivier Teboul, Félix de Chaumont Quitry, Marco TagliasacchiICLR 2021 · 181 citations
- Approximation and Learning with Deep Convolutional Models: a Kernel PerspectiveAlberto BiettiICLR 2022 · 33 citations
- ISNAS-DIP: Image-Specific Neural Architecture Search for Deep Image PriorMetin Ersin Arican, Ozgur Kara, Gustav Bredell, Ender KonukogluCVPR 2022 · 19 citations
