MathSpeech: Leveraging Small LMs for Accurate Conversion in Mathematical Speech-to-Formula
Sieun Hyeon, Kyudan Jung, Jaehee Won, Nam-Joon Kim, Hyun Gon Ryu, Hyuk-Jae Lee, Jaeyoung Do
Abstract
In various academic and professional settings, such as mathematics lectures or research presentations, it is often necessary to convey mathematical expressions orally. However, reading mathematical expressions aloud without accompanying visuals can significantly hinder comprehension, especially for those who are hearing-impaired or rely on subtitles due to language barriers. For instance, when a presenter reads Euler's Formula, current Automatic Speech Recognition (ASR) models often produce a verbose and error-prone textual description (e.g., e to the power of i x equals cosine of x plus i side of x), instead of the concise L A T E X format (i.e., e ix = cos(x) + i sin(x)), which hampers clear understanding and communication. To address this issue, we introduce MathSpeech, a novel pipeline that integrates ASR models with small Language Models (sLMs) to correct errors in mathematical expressions and accurately convert spoken expressions into structured L A T E X representations. Evaluated on a new dataset derived from lecture recordings, Math-Speech demonstrates L A T E X generation capabilities comparable to leading commercial Large Language Models (LLMs), while leveraging fine-tuned small language models of only 120M parameters. Specifically, in terms of CER, BLEU, and ROUGE scores for L A T E X translation, MathSpeech demonstrated significantly superior capabilities compared to GPT-4o. We observed a decrease in CER from 0.390 to 0.298, and higher ROUGE/BLEU scores compared to GPT-4o.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e33368e5-a6ea-4cfa-9a90-7778d27b6fd6Cited by top-tier papers1
Ask how each one uses itBuilds on5
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 1,267 citations
- Nougat: Neural Optical Understanding for Academic DocumentsLukas Blecher, Guillem Cucurull, Thomas Scialom, Robert StojnicICLR 2024 · 243 citations
- FastCorrect: Fast Error Correction with Edit Alignment for Automatic Speech RecognitionYichong Leng, Xu Tan, Linchen Zhu, Jin Xu et al.NeurIPS 2021 · 84 citations
- SoftCorrect: Error Correction with Soft Detection for Automatic Speech RecognitionYichong Leng, Xu Tan, Wenjie Liu, Kaitao Song et al.AAAI 2023 · 22 citations
Related papers
- Tree-Based Representation and Generation of Natural and Mathematical LanguageAlexander Scarlatos, Andrew S. LanACL 2023 · 13 citations
- GeoGPT4V: Towards Geometric Multi-modal Large Language Models with Geometric Image GenerationShihao Cai, Keqin Bao, Hangyu Guo, Jizhi Zhang et al.EMNLP 2024 · 10 citations
- Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong BaselineWeikang Bai, Yongkun Du, Yuchen Su, Yazhen Xie et al.AAAI 2026 · 2 citations
- MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMsZimu Lu, Aojun Zhou, Houxing Ren, Ke Wang et al.ACL 2024 · 11 citations
- MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical ReasoningKe Wang, Houxing Ren, Aojun Zhou, Zimu Lu et al.ICLR 2024 · 188 citations
