Dual-View Visual Contextualization for Web Navigation
Jihyung Kil, Chan Hee Song, Boyuan Zheng, Xiang Deng, Yu Su, Wei-Lun Chao
Abstract
Automatic web navigation aims to build a web agent that can follow language instructions to execute complex and diverse tasks on real-world websites. Existing work primarily takes HTML documents as input, which define the contents and action spaces (i.e., actionable elements and operations) of webpages. Nevertheless, HTML documents may not provide a clear task-related context for each element, making it hard to select the right (sequence of) actions. In this paper, we propose to contextualize HTML elements through their "dual views" in webpage screenshots: each HTML element has its corresponding bounding box and visual content in the screenshot. We build upon the insight-web developers tend to arrange task-related elements nearby on webpages to enhance user experiences-and propose to contextualize each element with its neighbor elements, using both textual and visual features. The resulting representations of HTML elements are more informative for the agent to take action. We validate our method on the recently released Mind2Web dataset, which features diverse navigation domains and tasks on real-world websites. Our method consistently outperforms the baseline in all the scenarios, including cross-task, cross-website, and cross-domain ones.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Web-Shepherd: Advancing PRMs for Reinforcing Web AgentsHyungjoo Chae, Sunghwan Kim, Junhee Cho, Seungone Kim et al.NeurIPS 2025 · 37 citations
- Attention-Driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models Without Fine-TuningHai-Ming Xu, Qi Chen, Lei Wang, Lingqiao LiuAAAI 2025 · 13 citations
- Watch and Learn: Learning to Use Computers from Online VideosChan Hee Song, Yiwen Song, Palash Goyal, Yu Su et al.CVPR 2026 · 8 citations
- Branch-and-Browse: Efficient and Controllable Web Exploration with Tree-Structured Reasoning and Action MemoryShiqi He, Yue Cui, Xinyu Ma, Yaliang Li et al.ACL 2026 · 5 citations
- ProactiveVA: Proactive Visual Analytics with LLM-Based UI AgentYuheng Zhao, Xueli Shu, Liwen Fan, Lin Gao et al.IEEE VIS 2025 · 4 citations
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
Related papers
- WEPO: Web Element Preference Optimization for LLM-based Web NavigationJiarun Liu, Jia Hao, Chunhong Zhang, Zheng HuAAAI 2025
- WebVLN: Vision-and-Language Navigation on WebsitesQi Chen, Dileepa Pitawela, Chongyang Zhao, Gengze Zhou et al.AAAI 2024 · 22 citations
- Harnessing Webpage UIs for Text-Rich Visual UnderstandingJunpeng Liu, Tianyue Ou, Yifan Song, Yuxiao Qu et al.ICLR 2025
- Lightweight Adaptive Topological Layout and Semantic Mapping in Vision-and-Language Navigation on WebsitesPingrui Lai, Zihao Xie, Hua YangAAAI 2026
- On the Multi-turn Instruction Following for Conversational Web AgentsYang Deng, Xuan Zhang, Wenxuan Zhang, Yifei Yuan et al.ACL 2024 · 5 citations
