ReactGenie: A Development Framework for Complex Multimodal Interactions Using Large Language Models
Jackie (Junrui) Yang, Yingtian Shi, Yuhan Zhang, Karina Li, Daniel Wan Rosli, Anisha Jain, Shuning Zhang, Tianshi Li, James A. Landay, Monica S. Lam
Abstract
By combining voice and touch interactions, multimodal interfaces can surpass the efficiency of either modality alone. Traditional multimodal frameworks require laborious developer work to support rich multimodal commands where the user's multimodal command involves possibly exponential combinations of actions/function invocations. This paper presents ReactGenie, a programming framework that better separates multimodal input from the computational model to enable developers to create efficient and capable multimodal interfaces with ease. ReactGenie translates multimodal user commands into NLPL (Natural Language Programming Language), a programming language we created, using a neural semantic parser based on large-language models. The ReactGenie runtime interprets the parsed NLPL and composes primitives in the computational model to implement complex user commands. As a result, React-Genie allows easy implementation and unprecedented richness in commands for end-users of multimodal apps. Our evaluation showed that 12 developers can learn and build a non-trivial Re-actGenie application in under 2.5 hours on average. In addition, compared with a traditional GUI, end-users can complete tasks faster and with less task load using ReactGenie apps.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec86805c-44d0-456c-b4c4-e9ecf3901059Cited by top-tier papers3
- How CO2STLY Is CHI? The Carbon Footprint of Generative AI in HCI Research and What We Should Do About ItNanna Inie, Jeanette Falk, Raghavendra SelvanCHI 2025 · 33 citations
- Classroom Simulacra: Building Contextual Student Generative Agents in Online Education for Learning Behavioral SimulationSonglin Xu, Hao-Ning Wen, Hongyi Pan, Dallas Dominguez et al.CHI 2025 · 13 citations
- GenieWizard: Multimodal App Feature Discovery with Large Language ModelsJackie (Junrui) Yang, Yingtian Shi, Chris Gu, Zhang Zheng et al.CHI 2025 · 6 citations
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Enabling Conversational Interaction with Mobile UI using Large Language ModelsBryan Wang, Gang Li, Yang LiCHI 2023 · 149 citations
- DoThisHere: Multimodal Interaction to Improve Cross-Application Tasks on Mobile DevicesJackie (Junrui) Yang, Monica S. Lam, James A. LandayUIST 2020 · 57 citations
- HybridTrak: Adding Full-Body Tracking to VR Using an Off-the-Shelf WebcamJackie (Junrui) Yang, Tuochao Chen, Fang Qin, Monica S. Lam et al.CHI 2022 · 39 citations
- Geno: A Developer Tool for Authoring Multimodal Interaction on Existing Web ApplicationsRitam Jyoti Sarmah, Yunpeng Ding, Di Wang, Cheuk Yin Phipson Lee et al.UIST 2020 · 19 citations
Related papers
- AutoM3L: An Automated Multimodal Machine Learning Framework with Large Language ModelsDaqin Luo, Chengjian Feng, Yuxuan Nong, Yiqing ShenACM MM 2024 · 16 citations
- From Navigation to Intention: Reframing the Web Experience through Goal-Driven InterfacesLuca Cordioli, Maristella MateraWWW 2026
- DeclarUI: Bridging Design and Development with Automated Declarative UI Code GenerationTing Zhou, Yanjie Zhao, Xinyi Hou, Xiaoyu Sun et al.FSE 2025 · 13 citations
- Controllable and Reliable Knowledge-Intensive Task-Oriented Conversational Agents with Declarative Genie WorksheetsHarshit Joshi, Shicheng Liu, James Chen, Larsen Weigle et al.ACL 2025
- APPL: A Prompt Programming Language for Harmonious Integration of Programs and Large Language Model PromptsHonghua Dong, Qidong Su, Yubo Gao, Zhaoyu Li et al.ACL 2025
