Geno: A Developer Tool for Authoring Multimodal Interaction on Existing Web Applications
Ritam Jyoti Sarmah, Yunpeng Ding, Di Wang, Cheuk Yin Phipson Lee, Toby Jia-Jun Li, Xiang 'Anthony' Chen
Abstract
Supporting voice commands in applications presents significant benefits to users. However, adding such support to existing GUI-based web apps is effort-consuming with a high learning barrier, as shown in our formative study, due to the lack of unified support for creating multimodal interfaces. We present Geno-a developer tool for adding the voice input modality to existing web apps without requiring significant NLP expertise. Geno provides a high-level workflow for developers to specify functionalities to be supported by voice (intents), create language models for detecting intents and the relevant information (parameters) from user utterances, and fulfill the intents by either programmatically invoking the corresponding functions or replaying GUI actions on the web app. Geno further supports multimodal references to GUI context in voice commands (e.g., "move this [event] to next week" while pointing at an event with the cursor). In a study, developers with little NLP expertise were able to add multimodal voice command support for two existing web apps using Geno.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 975ff000-5417-4e44-b4ce-a47bce9c50c4Cited by top-tier papers6
- Multi-Modal Repairs of Conversational Breakdowns in Task-Oriented DialogsToby Jia-Jun Li, Jingya Chen, Haijun Xia, Tom M. Mitchell et al.UIST 2020 · 98 citations
- Firefox Voice: An Open and Extensible Voice Assistant Built Upon the WebJulia Cambre, Alex C. Williams, Afsaneh Razi, Ian Bicking et al.CHI 2021 · 38 citations
- SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision ViewersZheng Ning, Brianna L. Wimer, Kaiwen Jiang, Keyi Chen et al.CHI 2024 · 27 citations
- ReactGenie: A Development Framework for Complex Multimodal Interactions Using Large Language ModelsJackie (Junrui) Yang, Yingtian Shi, Yuhan Zhang, Karina Li et al.CHI 2024 · 17 citations
- Voice and Touch Based Error-tolerant Multimodal Text Editing and Correction for SmartphonesMaozheng Zhao, Wenzhe Cui, I. V. Ramakrishnan, Shumin Zhai et al.UIST 2021 · 17 citations
Related papers
- GenieWizard: Multimodal App Feature Discovery with Large Language ModelsJackie (Junrui) Yang, Yingtian Shi, Chris Gu, Zhang Zheng et al.CHI 2025 · 6 citations
- Make-A-Voice: Revisiting Voice Large Language Models as Scalable Multilingual and Multitask LearnersRongjie Huang, Chunlei Zhang, Yongqi Wang, Dongchao Yang et al.ACL 2024 · 6 citations
- MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal InteractionsRamaneswaran Selvakumar, Ashish Seth, Nishit Anand, Utkarsh Tyagi et al.EMNLP 2025
- Discovering the Syntax and Strategies of Natural Language Programming with Generative Language ModelsEllen Jiang, Edwin Toh, Alejandra Molina, Kristen Olson et al.CHI 2022 · 76 citations
- Integrating Gaze and Speech for Enabling Implicit InteractionsAnam Ahmad Khan, Joshua Newn, James Bailey, Eduardo VellosoCHI 2022 · 16 citations
