Mapping Natural Language Instructions to Mobile UI Action Sequences
Yang Li, Jiacong He, Xin Zhou, Yuan Zhang, Jason Baldridge
摘要
We present a new problem: grounding natural language instructions to mobile user interface actions, and create three new datasets for it. For full task evaluation, we create PIX-ELHELP, a corpus that pairs English instructions with actions performed by people on a mobile UI emulator. To scale training, we decouple the language and action data by (a) annotating action phrase spans in HowTo instructions and (b) synthesizing grounded descriptions of actions for mobile user interfaces. We use a Transformer to extract action phrase tuples from long-range natural language instructions. A grounding Transformer then contextually represents UI objects using both their content and screen position and connects them to object descriptions. Given a starting screen and instruction, our model achieves 70.59% accuracy on predicting complete ground-truth action sequences in PIXELHELP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper93
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu 等ICML 2023 · 被引用 426 次
- A Real-World WebAgent with Planning, Long Context Understanding, and Program SynthesisIzzeddin Gur, Hiroki Furuta, Austin V. Huang, Mustafa Safdari 等ICLR 2024 · 被引用 359 次
- WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji 等ICML 2024 · 被引用 188 次
- Learnable Fourier Features for Multi-dimensional Spatial Positional EncodingYang Li, Si Si, Gang Li, Cho-Jui Hsieh 等NeurIPS 2021 · 被引用 171 次
相关 Paper
- Widget Captioning: Generating Natural Language Description for Mobile User Interface ElementsYang Li, Gang Li, Luheng He, Jingjie Zheng 等EMNLP 2020 · 被引用 46 次
- From Pixels to UI Actions: Learning to Follow Instructions via Graphical User InterfacesPeter Shaw, Mandar Joshi, James Cohan, Jonathan Berant 等NeurIPS 2023 · 被引用 89 次
- HelpViz: Automatic Generation of Contextual Visual Mobile Tutorials from Text-Based InstructionsMingyuan Zhong, Gang Li, Peggy Chi, Yang LiUIST 2021 · 被引用 27 次
- RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World UsersSuyu Ye, Haojun Shi, Darren Shih, Hyokun Yun 等AAAI 2026 · 被引用 17 次
- Instruction Induction: From Few Examples to Natural Language Task DescriptionsOr Honovich, Uri Shaham, Samuel R. Bowman, Omer LevyACL 2023 · 被引用 48 次
