Toward Interactive Dictation
Belinda Z. Li, Jason Eisner, Adam Pauls, Sam Thomson
摘要
Voice dictation is an increasingly important text input modality. Existing systems that allow both dictation and editing-by-voice restrict their command language to flat templates invoked by trigger words. In this work, we study the feasibility of allowing users to interrupt their dictation with spoken editing commands in open-ended natural language. We introduce a new task and dataset, TERTiUS, to experiment with such systems. To support this flexibility in real-time, a system must incrementally segment and classify spans of speech as either dictation or command, and interpret the spans that are commands. We experiment with using large pre-trained language models to predict the edited text, or alternatively, to predict a small text-editing program. Experiments show a natural trade-off between model accuracy and latency: a smaller model achieves 28% singlecommand interpretation accuracy with 1.3 seconds of latency, while a larger model achieves 55% with 7 seconds of latency. * Work performed during a research internship at Microsoft Semantic Machines. Just wanted to ask about the event on Friday the 23rd. Is the event still on? Just wanted to ask about the event on the 23rd, on Friday the 23rd. Is the event still on? Change"the event" to "it" in the last sentence. Just wanted to ask about the event on the 23rd. Just wanted to ask about the event on Friday the 23rd. Just wanted to check in about the event on Friday the 23rd. Is it still on?
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- CoditT5: Pretraining for Source Code and Natural Language EditingJiyang Zhang, Sheena Panthaplackel, Pengyu Nie, Junyi Jessy Li 等ASE 2022 · 被引用 81 次
- Online Semantic Parsing for Latency Reduction in Task-Oriented DialogueJiawei Zhou, Jason Eisner, Michael Newman, Emmanouil Antonios Platanios 等ACL 2022 · 被引用 6 次
相关 Paper
- A Full-duplex Speech Dialogue Scheme Based On Large Language ModelPeng Wang, Songshuo Lu, Yaohua Tang, Sijie Yan 等NeurIPS 2024
- Just Speak It: Minimize Cognitive Load for Eyes-Free Text Editing with a Smart Voice AssistantJiayue Fan, Chenning Xu, Chun Yu, Yuanchun ShiUIST 2021 · 被引用 17 次
- From User Perceptions to Technical Improvement: Enabling People Who Stutter to Better Use Speech RecognitionColin Lea, Zifang Huang, Jaya Narain, Lauren Tooley 等CHI 2023 · 被引用 35 次
- VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language ModelZuwei Long, Yunhang Shen, Chaoyou Fu, Heting Gao 等NeurIPS 2025 · 被引用 6 次
- Best of Both Worlds: Making High Accuracy Non-incremental Transformer-based Disfluency Detection IncrementalMorteza Rohanian, Julian HoughACL 2021
