Contrastive Learning from Synthetic Audio Doppelgängers
Manuel Cherep, Nikhil Singh
摘要
Learning robust audio representations currently demands extensive datasets of real-world sound recordings. By applying artificial transformations to these recordings, models can learn to recognize similarities despite subtle variations through techniques like contrastive learning. However, these transformations are only approximations of the true diversity found in real-world sounds, which are generated by complex interactions of physical processes, from vocal cord vibrations to the resonance of musical instruments. We propose a solution to both the data scale and transformation limitations, leveraging synthetic audio. By randomly perturbing the parameters of a sound synthesizer, we generate audio doppelgängers-synthetic positive pairs with causally manipulated variations in timbre, pitch, and temporal envelopes. These variations, difficult to achieve through augmentations of existing audio, provide a rich source of contrastive information. Despite the shift to randomly generated synthetic data, our method produces strong representations, outperforming real data on several standard audio classification tasks. Notably, our approach is lightweight, requires no data storage, and has only a single hyperparameter, which we extensively analyze. We offer this method as a complement to existing strategies for contrastive learning in audio, using synthesized sounds to reduce the data burden on practitioners.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- Effective Data Augmentation With Diffusion ModelsBrandon Trabucco, Kyle Doherty, Max Gurinas, Ruslan SalakhutdinovICLR 2024 · 被引用 380 次
- Generating Training Data with Language Models: Towards Zero-Shot Language UnderstandingYu Meng, Jiaxin Huang, Yu Zhang, Jiawei HanNeurIPS 2022 · 被引用 309 次
- StableRep: Synthetic Images from Text-to-Image Models Make Strong Visual Representation LearnersYonglong Tian, Lijie Fan, Phillip Isola, Huiwen Chang 等NeurIPS 2023 · 被引用 251 次
- Generative Models as a Data Source for Multiview Representation LearningAli Jahanian, Xavier Puig, Yonglong Tian, Phillip IsolaICLR 2022 · 被引用 148 次
相关 Paper
- Synthetic Data Can Also Teach: Synthesizing Effective Data for Unsupervised Visual Representation LearningYawen Wu, Zhepeng Wang, Dewen Zeng, Yiyu Shi 等AAAI 2023 · 被引用 20 次
- AudioMosaic: Contrastive Masked Audio Representation LearningHanxun Huang, Qizhou Wang, Xingjun Ma, Cihang Xie 等ICML 2026 · 被引用 2 次
- Synthio: Augmenting Small-Scale Audio Classification Datasets with Synthetic DataSreyan Ghosh, Sonal Kumar, Zhifeng Kong, Rafael Valle 等ICLR 2025
- Looking Similar, Sounding Different: Leveraging Counterfactual Cross-Modal Pairs for Audiovisual Representation LearningNikhil Singh, Chih-Wei Wu, Iroro Orife, Mahdi M. KalayehCVPR 2024
- Robust Dataset Condensation using Supervised Contrastive LearningNicole Hee-Yeon Kim, Hwanjun SongICCV 2025
