Dual-View Visual Contextualization for Web Navigation
Jihyung Kil, Chan Hee Song, Boyuan Zheng, Xiang Deng, Yu Su, Wei-Lun Chao
摘要
Automatic web navigation aims to build a web agent that can follow language instructions to execute complex and diverse tasks on real-world websites. Existing work primarily takes HTML documents as input, which define the contents and action spaces (i.e., actionable elements and operations) of webpages. Nevertheless, HTML documents may not provide a clear task-related context for each element, making it hard to select the right (sequence of) actions. In this paper, we propose to contextualize HTML elements through their "dual views" in webpage screenshots: each HTML element has its corresponding bounding box and visual content in the screenshot. We build upon the insight-web developers tend to arrange task-related elements nearby on webpages to enhance user experiences-and propose to contextualize each element with its neighbor elements, using both textual and visual features. The resulting representations of HTML elements are more informative for the agent to take action. We validate our method on the recently released Mind2Web dataset, which features diverse navigation domains and tasks on real-world websites. Our method consistently outperforms the baseline in all the scenarios, including cross-task, cross-website, and cross-domain ones.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Web-Shepherd: Advancing PRMs for Reinforcing Web AgentsHyungjoo Chae, Sunghwan Kim, Junhee Cho, Seungone Kim 等NeurIPS 2025 · 被引用 37 次
- Attention-Driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models Without Fine-TuningHai-Ming Xu, Qi Chen, Lei Wang, Lingqiao LiuAAAI 2025 · 被引用 13 次
- Watch and Learn: Learning to Use Computers from Online VideosChan Hee Song, Yiwen Song, Palash Goyal, Yu Su 等CVPR 2026 · 被引用 8 次
- Branch-and-Browse: Efficient and Controllable Web Exploration with Tree-Structured Reasoning and Action MemoryShiqi He, Yue Cui, Xinyu Ma, Yaliang Li 等ACL 2026 · 被引用 5 次
- ProactiveVA: Proactive Visual Analytics with LLM-Based UI AgentYuheng Zhao, Xueli Shu, Liwen Fan, Lin Gao 等IEEE VIS 2025 · 被引用 4 次
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
相关 Paper
- WEPO: Web Element Preference Optimization for LLM-based Web NavigationJiarun Liu, Jia Hao, Chunhong Zhang, Zheng HuAAAI 2025
- WebVLN: Vision-and-Language Navigation on WebsitesQi Chen, Dileepa Pitawela, Chongyang Zhao, Gengze Zhou 等AAAI 2024 · 被引用 22 次
- Harnessing Webpage UIs for Text-Rich Visual UnderstandingJunpeng Liu, Tianyue Ou, Yifan Song, Yuxiao Qu 等ICLR 2025
- Lightweight Adaptive Topological Layout and Semantic Mapping in Vision-and-Language Navigation on WebsitesPingrui Lai, Zihao Xie, Hua YangAAAI 2026
- On the Multi-turn Instruction Following for Conversational Web AgentsYang Deng, Xuan Zhang, Wenxuan Zhang, Yifei Yuan 等ACL 2024 · 被引用 5 次
