ReactGenie: A Development Framework for Complex Multimodal Interactions Using Large Language Models
Jackie (Junrui) Yang, Yingtian Shi, Yuhan Zhang, Karina Li, Daniel Wan Rosli, Anisha Jain, Shuning Zhang, Tianshi Li, James A. Landay, Monica S. Lam
摘要
By combining voice and touch interactions, multimodal interfaces can surpass the efficiency of either modality alone. Traditional multimodal frameworks require laborious developer work to support rich multimodal commands where the user's multimodal command involves possibly exponential combinations of actions/function invocations. This paper presents ReactGenie, a programming framework that better separates multimodal input from the computational model to enable developers to create efficient and capable multimodal interfaces with ease. ReactGenie translates multimodal user commands into NLPL (Natural Language Programming Language), a programming language we created, using a neural semantic parser based on large-language models. The ReactGenie runtime interprets the parsed NLPL and composes primitives in the computational model to implement complex user commands. As a result, React-Genie allows easy implementation and unprecedented richness in commands for end-users of multimodal apps. Our evaluation showed that 12 developers can learn and build a non-trivial Re-actGenie application in under 2.5 hours on average. In addition, compared with a traditional GUI, end-users can complete tasks faster and with less task load using ReactGenie apps.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- How CO2STLY Is CHI? The Carbon Footprint of Generative AI in HCI Research and What We Should Do About ItNanna Inie, Jeanette Falk, Raghavendra SelvanCHI 2025 · 被引用 33 次
- Classroom Simulacra: Building Contextual Student Generative Agents in Online Education for Learning Behavioral SimulationSonglin Xu, Hao-Ning Wen, Hongyi Pan, Dallas Dominguez 等CHI 2025 · 被引用 13 次
- GenieWizard: Multimodal App Feature Discovery with Large Language ModelsJackie (Junrui) Yang, Yingtian Shi, Chris Gu, Zhang Zheng 等CHI 2025 · 被引用 6 次
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Enabling Conversational Interaction with Mobile UI using Large Language ModelsBryan Wang, Gang Li, Yang LiCHI 2023 · 被引用 149 次
- DoThisHere: Multimodal Interaction to Improve Cross-Application Tasks on Mobile DevicesJackie (Junrui) Yang, Monica S. Lam, James A. LandayUIST 2020 · 被引用 57 次
- HybridTrak: Adding Full-Body Tracking to VR Using an Off-the-Shelf WebcamJackie (Junrui) Yang, Tuochao Chen, Fang Qin, Monica S. Lam 等CHI 2022 · 被引用 39 次
- Geno: A Developer Tool for Authoring Multimodal Interaction on Existing Web ApplicationsRitam Jyoti Sarmah, Yunpeng Ding, Di Wang, Cheuk Yin Phipson Lee 等UIST 2020 · 被引用 19 次
相关 Paper
- AutoM3L: An Automated Multimodal Machine Learning Framework with Large Language ModelsDaqin Luo, Chengjian Feng, Yuxuan Nong, Yiqing ShenACM MM 2024 · 被引用 16 次
- From Navigation to Intention: Reframing the Web Experience through Goal-Driven InterfacesLuca Cordioli, Maristella MateraWWW 2026
- DeclarUI: Bridging Design and Development with Automated Declarative UI Code GenerationTing Zhou, Yanjie Zhao, Xinyi Hou, Xiaoyu Sun 等FSE 2025 · 被引用 13 次
- Controllable and Reliable Knowledge-Intensive Task-Oriented Conversational Agents with Declarative Genie WorksheetsHarshit Joshi, Shicheng Liu, James Chen, Larsen Weigle 等ACL 2025
- APPL: A Prompt Programming Language for Harmonious Integration of Programs and Large Language Model PromptsHonghua Dong, Qidong Su, Yubo Gao, Zhaoyu Li 等ACL 2025
