DITTO: Data-efficient and Fair Targeted Subset Selection for ASR Accent Adaptation
Suraj Kothawade, Anmol Reddy Mekala, D. Chandra Sekhara Hetha Havya, Mayank Kothyari, Rishabh K. Iyer, Ganesh Ramakrishnan, Preethi Jyothi
摘要
State-of-the-art Automatic Speech Recognition (ASR) systems are known to exhibit disparate performance on varying speech accents. To improve performance on a specific target accent, a commonly adopted solution is to finetune the ASR model using accent-specific labeled speech. However, acquiring large amounts of labeled speech for specific target accents is challenging. Choosing an informative subset of speech samples that are most representative of the target accents becomes important for effective ASR finetuning. To address this problem, we propose DITTO (Data-efficient and faIr Targeted subseT selectiOn) that uses Submodular Mutual Information (SMI) functions as acquisition functions to find the most informative set of utterances matching a target accent within a fixed budget. An important feature of DITTO is that it supports fair targeting for multiple accents, i.e. it can automatically select representative data points from multiple accents when the ASR model needs to perform well on more than one accent. We show that DITTO is 3-5 times more label-efficient than other speech selection methods on the Indic-TTS and L2 datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- SIMILAR: Submodular Information Measures Based Active Learning In Realistic ScenariosSuraj Kothawade, Nathan Beck, KrishnaTeja Killamsetty, Rishabh K. IyerNeurIPS 2021 · 被引用 138 次
- PRISM: A Rich Class of Parameterized Submodular Information Measures for Guided Data Subset SelectionSuraj Kothawade, Vishal Kaushal, Ganesh Ramakrishnan, Jeff A. Bilmes 等AAAI 2022 · 被引用 66 次
- The Online Submodular Cover ProblemAnupam Gupta, Roie LevinSODA 2020 · 被引用 20 次
相关 Paper
- FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information GainRohan Deb, Kiran Koshy Thekumparampil, Kousha Kalantari, Gaurush Hiranandani 等ICML 2025
- Aligning Language Models with Demonstrated FeedbackOmar Shaikh, Michelle S. Lam, Joey Hejna, Yijia Shao 等ICLR 2025
- Accented Speech Recognition With Accent-specific CodebooksDarshan Prabhu, Preethi Jyothi, Sriram Ganapathy, Vinit UnniEMNLP 2023 · 被引用 5 次
- Adversarial Meta Sampling for Multilingual Low-Resource Speech RecognitionYubei Xiao, Ke Gong, Pan Zhou, Guolin Zheng 等AAAI 2021 · 被引用 37 次
- How Accents Confound: Probing for Accent Information in End-to-End Speech Recognition SystemsArchiki Prasad, Preethi JyothiACL 2020 · 被引用 18 次
