SummAct: Uncovering User Intentions Through Interactive Behaviour Summarisation
Guanhua Zhang, Mohamed Adel Naguib Ahmed, Zhiming Hu, Andreas Bulling
摘要
Recent work has highlighted the potential of modelling interactive behaviour analogously to natural language. We propose interactive behaviour summarisation as a novel computational task and demonstrate its usefulness for automatically uncovering latent user intentions while interacting with graphical user interfaces. To tackle this task, we introduce SummAct -a novel hierarchical method to summarise low-level input actions into high-level intentions. SummAct first identifies sub-goals from user actions using a large language model and in-context learning. High-level intentions are then obtained by fine-tuning the model using a novel UI element attention to preserve detailed context information embedded within UI elements during summarisation. Through a series of evaluations, we demonstrate that SummAct significantly outperforms baselines across desktop and mobile interfaces as well as interactive tasks by up to 21.9%. We further show three exciting interactive applications benefited from SummAct: interactive behaviour forecasting, automatic behaviour synonym identification, and language-based behaviour retrieval.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ProactiveVA: Proactive Visual Analytics with LLM-Based UI AgentYuheng Zhao, Xueli Shu, Liwen Fan, Lin Gao 等IEEE VIS 2025 · 被引用 4 次
- Codesigning Ripplet: an LLM-Assisted Assessment Authoring System Grounded in a Conceptual Model of Teachers' WorkflowsYuan Cui, Annabel Marie Goldman, Jovy Zhou, Xiaolin Liu 等CHI 2026 · 被引用 1 次
- GazeInterpreter: Parsing Eye Gaze to Generate Eye-Body-Coordinated NarrationsQing Chang, Zhiming HuAAAI 2026
- Evaluating Replay Techniques for Asynchronous Task Handover in Immersive AnalyticsZhengtai Gou, Junxiao Long, Tao Lu, Jian Zhao 等IEEE VR 2026
- Small Models, Big Results: Achieving Superior Intent Extraction through DecompositionDanielle Cohen, Yoni Halpern, Noam Kahlon, Joel Oren 等EMNLP 2025
它引用的顶会 Paper22
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric TasksWenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu 等NeurIPS 2023 · 被引用 725 次
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun 等ICML 2024 · 被引用 496 次
- The Pitfalls of Next-Token PredictionGregor Bachmann, Vaishnavh NagarajanICML 2024 · 被引用 163 次
- Enabling Conversational Interaction with Mobile UI using Large Language ModelsBryan Wang, Gang Li, Yang LiCHI 2023 · 被引用 149 次
相关 Paper
- Automatic Macro Mining from Interaction Traces at ScaleForrest Huang, Gang Li, Tao Li, Yang LiCHI 2024 · 被引用 11 次
- Screen2Words: Automatic Mobile UI Summarization with Multimodal LearningBryan Wang, Gang Li, Xin Zhou, Zhourong Chen 等UIST 2021 · 被引用 97 次
- UI-CTX: Understanding UI Behaviors with Code Contexts for Mobile ApplicationsJiawei Li, Jiahao Liu, Jian Mao, Jun Zeng 等NDSS 2025
- From Operation to Cognition: Automatic Modeling Cognitive Dependencies from User Demonstrations for GUI Task AutomationYiwen Yin, Yu Mei, Chun Yu, Toby Jia-Jun Li 等CHI 2025 · 被引用 8 次
- GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI TasksSaelyne Yang, Jaesang Yu, Yi-Hao Peng, Kevin Qinghong Lin 等CVPR 2026 · 被引用 5 次
