Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms
Zhangheng Li, Keen You, Haotian Zhang, Di Feng, Harsh Agrawal, Xiujun Li, Mohana Prasad Sathya Moorthy, Jeffrey Nichols, Yinfei Yang, Zhe Gan
Abstract
Building a generalist model for user interface (UI) understanding is challenging due to various foundational issues, such as platform diversity, resolution variation, and data limitation. In this paper, we introduce Ferret-UI 2, a multimodal large language model (MLLM) designed for universal UI understanding across a wide range of platforms, including iPhone, Android, iPad, Webpage, and AppleTV. Building on the foundation of Ferret-UI, Ferret-UI 2 introduces three key innovations: support for multiple platform types, high-resolution perception through adaptive scaling, and advanced task training data generation powered by GPT-4o with set-of-mark visual prompting. These advancements enable Ferret-UI 2 to perform complex, user-centered interactions, making it highly versatile and adaptable for the expanding diversity of platform ecosystems. Extensive empirical experiments on referring, grounding, user-centric advanced tasks (comprising 9 subtasks 5 platforms), GUIDE next-action prediction dataset, and GUI-World multi-platform benchmark demonstrate that Ferret-UI 2 significantly outperforms Ferret-UI, and also shows strong cross-platform transfer capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMsErik A. Daxberger, Nina Wenzel, David Griffiths, Haiming Gang et al.ICCV 2025 · 10 citations
- SO-Bench: A Structural Output Evaluation of Multimodal LLMDi Feng, Kaixin Ma, Feng Nan, Haofeng Chen et al.CVPR 2026
- Don't Click That: Teaching Web Agents to Resist Deceptive InterfacesYilin Zhang, Yingkai Hua, Chunyu Wei, Xin Wang et al.ACL 2026
Builds on21
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 1,477 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- Ferret: Refer and Ground Anything Anywhere at Any GranularityHaoxuan You, Haotian Zhang, Zhe Gan, Xianzhi Du et al.ICLR 2024 · 515 citations
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun et al.ICML 2024 · 496 citations
- A Real-World WebAgent with Planning, Long Context Understanding, and Program SynthesisIzzeddin Gur, Hiroki Furuta, Austin V. Huang, Mustafa Safdari et al.ICLR 2024 · 359 citations
Related papers
- Harnessing Webpage UIs for Text-Rich Visual UnderstandingJunpeng Liu, Tianyue Ou, Yifan Song, Yuxiao Qu et al.ICLR 2025
- UIPro: Unleashing Superior Interaction Capability for GUI AgentsHongxin Li, Jingran Su, Jingfan Chen, Zheng Ju et al.ICCV 2025
- Grounding Multimodal Large Language Model in GUI WorldWeixian Lei, Difei Gao, Mike Zheng ShouICLR 2025
- Grounding Multimodal Large Language Models to the WorldZhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao et al.ICLR 2024 · 1,170 citations
- MP-GUI: Modality Perception with MLLMs for GUI UnderstandingZiwei Wang, Weizhi Chen, Leyang Yang, Sheng Zhou et al.CVPR 2025
