SCRAPL: Scattering Transform with Random Paths for Machine Learning
Christopher Mitcheltree, Vincent Lostanlen, Emmanouil Benetos, Mathieu Lagrange
摘要
The Euclidean distance between wavelet scattering transform coefficients (known as paths) provides informative gradients for perceptual quality assessment of deep inverse problems in computer vision, speech, and audio processing. However, these transforms are computationally expensive when employed as differentiable loss functions for stochastic gradient descent due to their numerous paths, which significantly limits their use in neural network training. Against this problem, we propose "Scattering transform with Random Paths for machine Learning" (SCRAPL): a stochastic optimization scheme for efficient evaluation of multivariable scattering transforms. We implement SCRAPL for the joint time–frequency scattering transform (JTFS) which demodulates spectrotemporal patterns at multiple scales and rates, allowing a fine characterization of intermittent auditory textures. We apply SCRAPL to differentiable digital signal processing (DDSP), specifically, unsupervised sound matching of a granular synthesizer and the Roland TR-808 drum machine. We also propose an initialization heuristic based on importance sampling, which adapts SCRAPL to the perceptual content of the dataset, improving neural network convergence and evaluation performance. We make our code and audio samples available and provide SCRAPL as a Python package.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- DDSP: Differentiable Digital Signal ProcessingJesse H. Engel, Lamtharn Hantrakul, Chenjie Gu, Adam RobertsICLR 2020 · 被引用 467 次
- Pruned Graph Scattering TransformsVassilis N. Ioannidis, Siheng Chen, Georgios B. GiannakisICLR 2020 · 被引用 28 次
- Parametric Scattering NetworksShanel Gauthier, Benjamin Thérien, Laurent Alsène-Racicot, Muawiz Chaudhary 等CVPR 2022 · 被引用 16 次
- Regularizing Neural Networks with Meta-Learning Generative ModelsShin'ya Yamaguchi, Daiki Chijiwa, Sekitoshi Kanai, Atsutoshi Kumagai 等NeurIPS 2023 · 被引用 10 次
- ADAM Optimization with Adaptive Batch SelectionGyu-Yeol Kim, Min-hwan OhICLR 2025
相关 Paper
- Learning Temporal Resolution in Spectrogram for Audio ClassificationHaohe Liu, Xubo Liu, Qiuqiang Kong, Wenwu Wang 等AAAI 2024 · 被引用 15 次
- DDSL: Deep Differentiable Simplex Layer for Learning Geometric SignalsChiyu Max Jiang, Dana Lynn Ona Lansigan, Philip Marcus, Matthias NießnerICCV 2019 · 被引用 12 次
- Differentiable Geometric Acoustic Path Tracing using Time-Resolved Path Replay BackpropagationUgo Paavo Finnendahl, Markus Worchel, Tobias Jüterbock, Daniel Wujecki 等SIGGRAPH 2025 · 被引用 3 次
- LSCD: Lomb-Scargle Conditioned Diffusion for Time series ImputationElizabeth Fons, Alejandro Sztrajman, Yousef El-Laham, Luciana Ferrer 等ICML 2025
- Grad-TTS: A Diffusion Probabilistic Model for Text-to-SpeechVadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova 等ICML 2021 · 被引用 715 次
