OmniActions: Predicting Digital Actions in Response to Real-World Multimodal Sensory Inputs with LLMs
Jiahao Nick Li, Yan Xu, Tovi Grossman, Stephanie Santosa, Michelle Li
摘要
The progression to “Pervasive Augmented Reality” envisions easy access to multimodal information continuously. However, in many everyday scenarios, users are occupied physically, cognitively or socially. This may increase the friction to act upon the multimodal information that users encounter in the world. To reduce such friction, future interactive interfaces should intelligently provide quick access to digital actions based on users’ context. To explore the range of possible digital actions, we conducted a diary study that required participants to capture and share the media that they intended to perform actions on (e.g., images or audio), along with their desired actions and other contextual information. Using this data, we generated a holistic design space of digital follow-up actions that could be performed in response to different types of multimodal sensory inputs. We then designed OmniActions, a pipeline powered by large language models (LLMs) that processes multimodal sensory inputs and predicts follow-up actions on the target information grounded in the derived design space. Using the empirical data collected in the diary study, we performed quantitative evaluations on three variations of LLM techniques (intent classification, in-context learning and finetuning) and identified the most effective technique for our task. Additionally, as an instantiation of the pipeline, we developed an interactive prototype and reported preliminary user feedback about how people perceive and react to the action predictions and its errors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Augmented Object Intelligence with XR-ObjectsMustafa Doga Dogan, Eric J. Gonzalez, Karan Ahuja, Ruofei Du 等UIST 2024 · 被引用 64 次
- Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature ReviewRock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas 等CHI 2025 · 被引用 51 次
- Satori 悟り: Towards Proactive AR Assistant with Belief-Desire-Intention User ModelingChenyi Li, Guande Wu, Gromit Yeuk-Yin Chan, Dishita G. Turakhia 等CHI 2025 · 被引用 49 次
- How CO2STLY Is CHI? The Carbon Footprint of Generative AI in HCI Research and What We Should Do About ItNanna Inie, Jeanette Falk, Raghavendra SelvanCHI 2025 · 被引用 33 次
- AiGet: Transforming Everyday Moments into Hidden Knowledge Discovery with AI Assistance on Smart GlassesRunze Cai, Nuwan Janaka, Hyeongcheol Kim, Yang Chen 等CHI 2025 · 被引用 26 次
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
相关 Paper
- HMotionGPT: Aligning Hand Motions and Natural Language for Activity Understanding with Smart RingsYang Gao, Dong She, Wolin Liang, Chiyue Wang 等UbiComp 2026
- Explainable XR: Understanding User Behaviors of XR Environments Using LLM-Assisted Analytics FrameworkYoonsang Kim, Zainab Aamir, Mithilesh Kumar Singh, Saeed Boorboor 等IEEE VR 2025 · 被引用 27 次
- RoboOmni: Proactive Robot Manipulation in Omni-modal ContextSiyin Wang, Jinlan Fu, Feihong Liu, Xinzhe He 等ICLR 2026 · 被引用 9 次
- Pro 2 Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural TasksLilin Xu, Bufang Yang, Siyang Jiang, Kaiwei Liu 等UbiComp 2026
- OmniQuery: Contextually Augmenting Captured Multimodal Memories to Enable Personal Question AnsweringJiahao Nick Li, Zhuohao Jerry Zhang, Jiaju MaCHI 2025 · 被引用 22 次
