Investigating the effect of auxiliary objectives for the automated grading of learner English speech transcriptions
Hannah Craighead, Andrew Caines, Paula Buttery, Helen Yannakoudakis
Abstract
We address the task of automatically grading the language proficiency of spontaneous speech based on textual features from automatic speech recognition transcripts. Motivated by recent advances in multi-task learning, we develop neural networks trained in a multi-task fashion that learn to predict the proficiency level of non-native English speakers by taking advantage of inductive transfer between the main task (grading) and auxiliary prediction tasks: morpho-syntactic labeling, language modeling, and native language identification (L1). We encode the transcriptions with both bi-directional recurrent neural networks and with bi-directional representations from transformers, compare against a featurerich baseline, and analyse performance at different proficiency levels and with transcriptions of varying error rates. Our best performance comes from a transformer encoder with L1 prediction as an auxiliary task. We discuss areas for improvement and potential applications for text-only speech scoring.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Related papers
- LCMA-SRT: Language-Conditional Mixture-of-Experts Adapters for Joint Multilingual Speech Recognition and TranslationNanjie Li, Xiaoyong Guo, Hao Huang, Haihua Xu et al.ACL 2026
- Improving Speech Translation by Understanding and Learning from the Auxiliary Text Translation TaskYun Tang, Juan Miguel Pino, Xian Li, Changhan Wang et al.ACL 2021
- GradTS: A Gradient-Based Automatic Auxiliary Task Selection Method Based on Transformer NetworksWeicheng Ma, Renze Lou, Kai Zhang, Lili Wang et al.EMNLP 2021 · 4 citations
- Read to Hear: A Zero-Shot Pronunciation Assessment Using Textual Descriptions and LLMsYu-Wen Chen, Melody Ma, Julia HirschbergEMNLP 2025
- Unified Speech-Text Pre-training for Speech Translation and RecognitionYun Tang, Hongyu Gong, Ning Dong, Changhan Wang et al.ACL 2022 · 104 citations
