Enhancing Textbooks with Visuals from the Web for Improved Learning
Janvijay Singh, Vilém Zouhar, Mrinmaya Sachan
摘要
Textbooks are one of the main mediums for delivering high-quality education to students. In particular, explanatory and illustrative visuals play a key role in retention, comprehension and general transfer of knowledge. However, many textbooks lack these interesting visuals to support student learning. In this paper, we investigate the effectiveness of vision-language models to automatically enhance textbooks with images from the web. We collect a dataset of e-textbooks in the math, science, social science and business domains. We then set up a text-image matching task that involves retrieving and appropriately assigning web images to textbooks, which we frame as a matching optimization problem. Through a crowd-sourced evaluation, we verify that (1) while the original textbook images are rated higher, automatically assigned ones are not far behind, and (2) the precise formulation of the optimization problem matters. We release the dataset of textbooks with an associated image bank to inspire further research in this intersectional area of computer vision and NLP for education.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper2
相关 Paper
- 2.5 Years in Class: A Multimodal Textbook for Vision-Language PretrainingWenqi Zhang, Hang Zhang, Xin Li, Jiashuo Sun 等ICCV 2025 · 被引用 2 次
- Learning the Visualness of Text Using Large Vision-Language ModelsGaurav Verma, Ryan A. Rossi, Christopher Tensmeyer, Jiuxiang Gu 等EMNLP 2023 · 被引用 2 次
- Lecture Presentations Multimodal Dataset: Towards Understanding Multimodality in Educational VideosDong Won Lee, Chaitanya Ahuja, Paul Pu Liang, Sanika Natu 等ICCV 2023 · 被引用 20 次
- VisualWebInstruct: Scaling up Multimodal Instruction Data through Web SearchYiming Jia, Jiachen Li, Xiang Yue, Bo Li 等EMNLP 2025 · 被引用 29 次
- MemeCap: A Dataset for Captioning and Interpreting MemesEunjeong Hwang, Vered ShwartzEMNLP 2023 · 被引用 14 次
