A Suite of Generative Tasks for Multi-Level Multimodal Webpage Understanding
Andrea Burns, Krishna Srinivasan, Joshua Ainslie, Geoff Brown, Bryan A. Plummer, Kate Saenko, Jianmo Ni, Mandy Guo
摘要
Webpages have been a rich, scalable resource for vision-language and language only tasks. Yet only pieces of webpages are kept in existing datasets: image-caption pairs, long text articles, or raw HTML, never all in one place. Webpage tasks have resultingly received little attention and structured image-text data left underused. To study multimodal webpage understanding, we introduce the Wikipedia Webpage suite (WikiWeb2M) containing 2M pages with all of the associated image, text, and structure data 1 . We verify its utility on three generative tasks: page description generation, section summarization, and contextual image captioning. We design a novel attention mechanism Prefix Global, which selects the most relevant image and text content as global tokens to attend to the rest of the webpage for context. By using page structure to separate such tokens, it performs better than full attention with lower computational complexity. Extensive experiments show that the new data in WikiWeb2M improves task performance compared to prior work.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- UniGraph2: Learning a Unified Embedding Space to Bind Multimodal GraphsYufei He, Yuan Sui, Xiaoxin He, Yue Liu 等WWW 2025 · 被引用 37 次
- OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal RetrievalWei Yang, Jingjing Fu, Rui Wang, Jinyu Wang 等ACL 2025 · 被引用 11 次
- IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?Yang Chen, Minghao Liu, Yufan Shen, Yunwen Li 等ICLR 2026 · 被引用 11 次
- Cross-Contrastive Clustering for Multimodal Attributed Graphs with Dual Graph FilteringHaoran Zheng, Renchi Yang, Hongtao Wang, Jianliang XuKDD 2026 · 被引用 7 次
- REAL: Resolving Knowledge Conflicts in Knowledge-Intensive Visual Question Answering via Reasoning-Pivot AlignmentKai Ye, Xianwei Mao, Sheng Zhou, Zirui Shao 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
相关 Paper
- Show, Interpret and Tell: Entity-Aware Contextualised Image Captioning in WikipediaKhanh Nguyen, Ali Furkan Biten, Andrés Mafla, Lluís Gómez 等AAAI 2023 · 被引用 15 次
- WebCode2M: A Real-World Dataset for Code Generation from Webpage DesignsYi Gui, Zhen Li, Yao Wan, Yemin Shi 等WWW 2025 · 被引用 38 次
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu 等ICML 2023 · 被引用 426 次
- UniMS: A Unified Framework for Multimodal Summarization with Knowledge DistillationZhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang 等AAAI 2022 · 被引用 61 次
- Cross-Lingual Phrase RetrievalHeqi Zheng, Xiao Zhang, Zewen Chi, Heyan Huang 等ACL 2022 · 被引用 1 次
