Just Speak It: Minimize Cognitive Load for Eyes-Free Text Editing with a Smart Voice Assistant
Jiayue Fan, Chenning Xu, Chun Yu, Yuanchun Shi
Abstract
Entering text precisely by voice, users might encounter colloquial inserts, inappropriate wording, and recognition errors, which brings difficulties to voice editing. Users need to locate the errors and then correct them. In eyes-free scenarios, this select-modify mode brings a cognitive burden and a risk of error. This paper introduces neural networks and pre-trained models to understand users’ revision intention based on semantics, reducing the need for the information from users’ statements. We present two strategies. One is to remove the colloquial inserts automatically. The other is to allow users to edit by just speaking out the target words without having to say the context and the incorrect text. Accordingly, our approach can predict whether to insert or replace, the incorrect text to replace, and the position to insert. We implement these strategies in SmartEdit, an eyes-free voice input agent controlled with earphone buttons. The evaluation shows that our techniques reduce the cognitive load and decrease the average failure rate by 54.1% compared to descriptive command or re-speaking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d6fffa4-a7ff-4152-8829-7d8782034cf6Cited by top-tier papers3
- Rambler: Supporting Writing With Speech via LLM-Assisted Gist ManipulationSusan Lin, Jeremy Warner, J. D. Zamfirescu-Pereira, Matthew G. Lee et al.CHI 2024 · 38 citations
- GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and LearningMinh Duc Vu, Han Wang, Jieshan Chen, Zhuang Li et al.UIST 2024 · 17 citations
- Typist Experiment: an Investigation of Human-to-Human Dictation via Role-play to Inform Voice-based Text AuthoringCan Liu, Siying Hu, Li Feng, Mingming FanCSCW 2022 · 7 citations
Builds on3
- Spelling Error Correction with Soft-Masked BERTShaohua Zhang, Haoran Huang, Jicong Liu, Hang LiACL 2020 · 204 citations
- EYEditor: Towards On-the-Go Heads-Up Text Editing Using Voice and Manual InputDebjyoti Ghosh, Pin Sym Foong, Shengdong Zhao, Can Liu et al.CHI 2020 · 42 citations
- Leveraging Error Correction in Voice-based Text Entry by Talk-and-GazeKorok Sengupta, Sabin Bhattarai, Sayan Sarcar, I. Scott MacKenzie et al.CHI 2020 · 14 citations
Related papers
- Platform for Studying Self-Repairing Auto-Corrections in Mobile Text Entry based on Brain Activity, Gaze, and ContextFelix Putze, Tilman Ihrig, Tanja Schultz, Wolfgang StuerzlingerCHI 2020 · 10 citations
- Voice and Touch Based Error-tolerant Multimodal Text Editing and Correction for SmartphonesMaozheng Zhao, Wenzhe Cui, I. V. Ramakrishnan, Shumin Zhai et al.UIST 2021 · 17 citations
- SmartFreeEdit: Mask-Free Spatial-Aware Image Editing with Complex Instruction UnderstandingQianqian Sun, Jixiang Luo, Dell Zhang, Xuelong LiACM MM 2025
- Toward Interactive DictationBelinda Z. Li, Jason Eisner, Adam Pauls, Sam ThomsonACL 2023 · 2 citations
- X2T: Training an X-to-Text Typing Interface with Online Learning from User FeedbackJensen Gao, Siddharth Reddy, Glen Berseth, Nicholas Hardy et al.ICLR 2021 · 10 citations
