SummAct: Uncovering User Intentions Through Interactive Behaviour Summarisation
Guanhua Zhang, Mohamed Adel Naguib Ahmed, Zhiming Hu, Andreas Bulling
Abstract
Recent work has highlighted the potential of modelling interactive behaviour analogously to natural language. We propose interactive behaviour summarisation as a novel computational task and demonstrate its usefulness for automatically uncovering latent user intentions while interacting with graphical user interfaces. To tackle this task, we introduce SummAct -a novel hierarchical method to summarise low-level input actions into high-level intentions. SummAct first identifies sub-goals from user actions using a large language model and in-context learning. High-level intentions are then obtained by fine-tuning the model using a novel UI element attention to preserve detailed context information embedded within UI elements during summarisation. Through a series of evaluations, we demonstrate that SummAct significantly outperforms baselines across desktop and mobile interfaces as well as interactive tasks by up to 21.9%. We further show three exciting interactive applications benefited from SummAct: interactive behaviour forecasting, automatic behaviour synonym identification, and language-based behaviour retrieval.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 33675166-8097-417b-af13-0479bb2adfdfCited by top-tier papers5
- ProactiveVA: Proactive Visual Analytics with LLM-Based UI AgentYuheng Zhao, Xueli Shu, Liwen Fan, Lin Gao et al.IEEE VIS 2025 · 4 citations
- Codesigning Ripplet: an LLM-Assisted Assessment Authoring System Grounded in a Conceptual Model of Teachers' WorkflowsYuan Cui, Annabel Marie Goldman, Jovy Zhou, Xiaolin Liu et al.CHI 2026 · 1 citation
- GazeInterpreter: Parsing Eye Gaze to Generate Eye-Body-Coordinated NarrationsQing Chang, Zhiming HuAAAI 2026
- Evaluating Replay Techniques for Asynchronous Task Handover in Immersive AnalyticsZhengtai Gou, Junxiao Long, Tao Lu, Jian Zhao et al.IEEE VR 2026
- Small Models, Big Results: Achieving Superior Intent Extraction through DecompositionDanielle Cohen, Yoni Halpern, Noam Kahlon, Joel Oren et al.EMNLP 2025
Builds on22
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric TasksWenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu et al.NeurIPS 2023 · 725 citations
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun et al.ICML 2024 · 496 citations
- The Pitfalls of Next-Token PredictionGregor Bachmann, Vaishnavh NagarajanICML 2024 · 163 citations
- Enabling Conversational Interaction with Mobile UI using Large Language ModelsBryan Wang, Gang Li, Yang LiCHI 2023 · 149 citations
Related papers
- Automatic Macro Mining from Interaction Traces at ScaleForrest Huang, Gang Li, Tao Li, Yang LiCHI 2024 · 11 citations
- Screen2Words: Automatic Mobile UI Summarization with Multimodal LearningBryan Wang, Gang Li, Xin Zhou, Zhourong Chen et al.UIST 2021 · 97 citations
- UI-CTX: Understanding UI Behaviors with Code Contexts for Mobile ApplicationsJiawei Li, Jiahao Liu, Jian Mao, Jun Zeng et al.NDSS 2025
- From Operation to Cognition: Automatic Modeling Cognitive Dependencies from User Demonstrations for GUI Task AutomationYiwen Yin, Yu Mei, Chun Yu, Toby Jia-Jun Li et al.CHI 2025 · 8 citations
- GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI TasksSaelyne Yang, Jaesang Yu, Yi-Hao Peng, Kevin Qinghong Lin et al.CVPR 2026 · 5 citations
