Toward Interactive Dictation
Belinda Z. Li, Jason Eisner, Adam Pauls, Sam Thomson
Abstract
Voice dictation is an increasingly important text input modality. Existing systems that allow both dictation and editing-by-voice restrict their command language to flat templates invoked by trigger words. In this work, we study the feasibility of allowing users to interrupt their dictation with spoken editing commands in open-ended natural language. We introduce a new task and dataset, TERTiUS, to experiment with such systems. To support this flexibility in real-time, a system must incrementally segment and classify spans of speech as either dictation or command, and interpret the spans that are commands. We experiment with using large pre-trained language models to predict the edited text, or alternatively, to predict a small text-editing program. Experiments show a natural trade-off between model accuracy and latency: a smaller model achieves 28% singlecommand interpretation accuracy with 1.3 seconds of latency, while a larger model achieves 55% with 7 seconds of latency. * Work performed during a research internship at Microsoft Semantic Machines. Just wanted to ask about the event on Friday the 23rd. Is the event still on? Just wanted to ask about the event on the 23rd, on Friday the 23rd. Is the event still on? Change"the event" to "it" in the last sentence. Just wanted to ask about the event on the 23rd. Just wanted to ask about the event on Friday the 23rd. Just wanted to check in about the event on Friday the 23rd. Is it still on?
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ce8f10cc-afde-4caf-8f71-eb620c6a3408Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- CoditT5: Pretraining for Source Code and Natural Language EditingJiyang Zhang, Sheena Panthaplackel, Pengyu Nie, Junyi Jessy Li et al.ASE 2022 · 81 citations
- Online Semantic Parsing for Latency Reduction in Task-Oriented DialogueJiawei Zhou, Jason Eisner, Michael Newman, Emmanouil Antonios Platanios et al.ACL 2022 · 6 citations
Related papers
- A Full-duplex Speech Dialogue Scheme Based On Large Language ModelPeng Wang, Songshuo Lu, Yaohua Tang, Sijie Yan et al.NeurIPS 2024
- Just Speak It: Minimize Cognitive Load for Eyes-Free Text Editing with a Smart Voice AssistantJiayue Fan, Chenning Xu, Chun Yu, Yuanchun ShiUIST 2021 · 17 citations
- From User Perceptions to Technical Improvement: Enabling People Who Stutter to Better Use Speech RecognitionColin Lea, Zifang Huang, Jaya Narain, Lauren Tooley et al.CHI 2023 · 35 citations
- VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language ModelZuwei Long, Yunhang Shen, Chaoyou Fu, Heting Gao et al.NeurIPS 2025 · 6 citations
- Best of Both Worlds: Making High Accuracy Non-incremental Transformer-based Disfluency Detection IncrementalMorteza Rohanian, Julian HoughACL 2021
