Creating General User Models from Computer Use
Omar Shaikh, Shardul Sapkota, Shan Rizvi, Eric Horvitz, Joon Sung Park, Diyi Yang, Michael S. Bernstein
摘要
Human-computer interaction has long imagined technology that understands us—from our preferences and habits, to the timing and purpose of our everyday actions. Yet current user models remain fragmented, narrowly tailored to specific applications, and incapable of the flexible, cross-context reasoning required to fulfill these visions. This paper presents an architecture for a general user model (GUM) that learns about you by observing any interaction you have with your computer. The GUM takes as input any unstructured observation of a user (e.g., device screenshots) and constructs confidence-weighted natural language propositions that capture that user’s behavior, knowledge, beliefs, and preferences. GUMs can infer that a user is preparing for a wedding they’re attending from a message thread with a friend. Or recognize that a user is struggling with a collaborator’s feedback on a draft paper by observing multiple stalled edits and a switch to reading related work. GUMs introduce an architecture that infers new propositions about a user from multimodal observations, retrieves related propositions for context, and continuously revises existing propositions. To illustrate the breadth of applications that GUMs enable, we demonstrate how they augment chat-based assistants with contextual understanding, manage OS notifications to surface important information only when needed, and enable interactive agents that adapt to user preferences across applications. We also instantiate a new class of proactive assistants (Gumbos) that discover and execute useful suggestions on a user’s behalf based on their GUM. In our evaluations, we find that GUMs make calibrated and accurate inferences about users, and that assistants built on GUMs proactively identify and perform actions of meaningful value that users wouldn’t think to request explicitly. Altogether, GUMs introduce new methods that leverage large multimodal models to understand unstructured user context, enabling both long-standing visions of HCI and entirely new interactive systems that anticipate user needs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and RectificationStefan Krsteski, Giuseppe Russo, Serina Chang, Robert West 等ACL 2026 · 被引用 10 次
- Cocoa: Co-Planning and Co-Execution with AI AgentsK. J. Kevin Feng, Kevin Pu, Matt Latzke, Tal August 等CHI 2026 · 被引用 6 次
- Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human OversightJingyu Tang, Chaoran Chen, Jiawen Li, Zhiping Zhang 等CHI 2026 · 被引用 4 次
- Just-In-Time Objectives: A General Approach for Specialized AI InteractionsMichelle S. Lam, Omar Shaikh, Hallie Xu, Alice Guo 等CHI 2026 · 被引用 3 次
- HistoryPalette: Supporting Exploration and Reuse of Past Alternatives in Image Generation and EditingKarim Benharrak, Amy PavelCHI 2026 · 被引用 2 次
它引用的顶会 Paper19
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 被引用 1,030 次
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 被引用 892 次
相关 Paper
- Sensible Agent: A Framework for Unobtrusive Interaction with Proactive AR AgentsGeonsun Lee, Min Xia, Nels Numan, Xun Qian 等UIST 2025 · 被引用 15 次
- Pro 2 Assist: Continuous Step-aware Proactive Assistance with Multi-modal Egocentric Perception for Long-horizon Procedural TasksLilin Xu, Bufang Yang, Siyang Jiang, Kaiwei Liu 等UbiComp 2026
- HMotionGPT: Aligning Hand Motions and Natural Language for Activity Understanding with Smart RingsYang Gao, Dong She, Wolin Liang, Chiyue Wang 等UbiComp 2026
- ProMemAssist: Exploring Timely Proactive Assistance Through Working Memory Modeling in Multi-Modal Wearable DevicesKevin Pu, Ting Zhang, Naveen Sendhilnathan, Sebastian Freitag 等UIST 2025 · 被引用 9 次
- AMMA: Adaptive Multimodal Assistants Through Automated State Tracking and User Model-Directed Guidance PlanningJackie (Junrui) Yang, Leping Qiu, Emmanuel Angel Corona-Moreno, Louisa Shi 等IEEE VR 2024 · 被引用 11 次
