Voice and Touch Based Error-tolerant Multimodal Text Editing and Correction for Smartphones
Maozheng Zhao, Wenzhe Cui, I. V. Ramakrishnan, Shumin Zhai, Xiaojun Bi
Abstract
Editing operations such as cut, copy, paste, and correcting errors in typed text are often tedious and challenging to perform on smartphones. In this paper, we present VT, a voice and touch-based multi-modal text editing and correction method for smartphones. To edit text with VT, the user glides over a text fragment with a finger and dictates a command, such as "bold" to change the format of the fragment, or the user can tap inside a text area and speak a command such as "highlight this paragraph" to edit the text. For text correcting, the user taps approximately at the area of erroneous text fragment and dictates the new content for substitution or insertion. VT combines touch and voice inputs with language context such as language model and phrase similarity to infer a user's editing intention, which can handle ambiguities and noisy input signals. It is a great advantage over the existing error correction methods (e.g., iOS's Voice Control) which require precise cursor control or text selection. Our evaluation shows that VT significantly improves the efficiency of text editing and text correcting on smartphones over the touch-only method and the iOS's Voice Control method. Our user studies showed that VT reduced the text editing time by 30.80%, and text correcting time by 29.97% over the touch-only method. VT reduced the text editing time by 30.81%, and text correcting time by 47.96% over the iOS's Voice Control method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Rambler: Supporting Writing With Speech via LLM-Assisted Gist ManipulationSusan Lin, Jeremy Warner, J. D. Zamfirescu-Pereira, Matthew G. Lee et al.CHI 2024 · 38 citations
- Typist Experiment: an Investigation of Human-to-Human Dictation via Role-play to Inform Voice-based Text AuthoringCan Liu, Siying Hu, Li Feng, Mingming FanCSCW 2022 · 7 citations
- SwivelTouch: Boosting Touchscreen Input with 3D Finger Rotation GestureChentao Li, Jinyang Yu, Ke He, Jianjiang Feng et al.UbiComp 2024 · 3 citations
Builds on4
- Multi-Modal Repairs of Conversational Breakdowns in Task-Oriented DialogsToby Jia-Jun Li, Jingya Chen, Haijun Xia, Tom M. Mitchell et al.UIST 2020 · 98 citations
- DoThisHere: Multimodal Interaction to Improve Cross-Application Tasks on Mobile DevicesJackie (Junrui) Yang, Monica S. Lam, James A. LandayUIST 2020 · 57 citations
- JustCorrect: Intelligent Post Hoc Text Correction Techniques on SmartphonesWenzhe Cui, Suwen Zhu, Mingrui Ray Zhang, H. Andrew Schwartz et al.UIST 2020 · 24 citations
- Geno: A Developer Tool for Authoring Multimodal Interaction on Existing Web ApplicationsRitam Jyoti Sarmah, Yunpeng Ding, Di Wang, Cheuk Yin Phipson Lee et al.UIST 2020 · 19 citations
Related papers
- Swap: A Replacement-based Text Revision Technique for Mobile DevicesYang Li, Sayan Sarcar, Sunjun Kim, Xiangshi RenCHI 2020 · 9 citations
- MMPE: A Multi-Modal Interface for Post-Editing Machine TranslationNico Herbig, Tim Düwel, Santanu Pal, Kalliopi Meladaki et al.ACL 2020 · 24 citations
- Tap&Say: Touch Location-Informed Large Language Model for Multimodal Text Correction on SmartphonesMaozheng Zhao, Michael Xuelin Huang, Nathan G. Huang, Shanqing Cai et al.CHI 2025 · 6 citations
- Just Speak It: Minimize Cognitive Load for Eyes-Free Text Editing with a Smart Voice AssistantJiayue Fan, Chenning Xu, Chun Yu, Yuanchun ShiUIST 2021 · 17 citations
- Enabling Auto-Correction on Soft Braille KeyboardDan Zhang, Yan Ma, Glenn Dausch, William H. Seiple et al.UIST 2025 · 1 citation
