Enhancing Textbooks with Visuals from the Web for Improved Learning
Janvijay Singh, Vilém Zouhar, Mrinmaya Sachan
Abstract
Textbooks are one of the main mediums for delivering high-quality education to students. In particular, explanatory and illustrative visuals play a key role in retention, comprehension and general transfer of knowledge. However, many textbooks lack these interesting visuals to support student learning. In this paper, we investigate the effectiveness of vision-language models to automatically enhance textbooks with images from the web. We collect a dataset of e-textbooks in the math, science, social science and business domains. We then set up a text-image matching task that involves retrieving and appropriately assigning web images to textbooks, which we frame as a matching optimization problem. Through a crowd-sourced evaluation, we verify that (1) while the original textbook images are rated higher, automatically assigned ones are not far behind, and (2) the precise formulation of the optimization problem matters. We release the dataset of textbooks with an associated image bank to inspire further research in this intersectional area of computer vision and NLP for education.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03505952-b90a-4d0d-a6a4-fcadc82088a3Cited by top-tier papers1
Ask how each one uses itBuilds on2
Related papers
- 2.5 Years in Class: A Multimodal Textbook for Vision-Language PretrainingWenqi Zhang, Hang Zhang, Xin Li, Jiashuo Sun et al.ICCV 2025 · 2 citations
- Learning the Visualness of Text Using Large Vision-Language ModelsGaurav Verma, Ryan A. Rossi, Christopher Tensmeyer, Jiuxiang Gu et al.EMNLP 2023 · 2 citations
- Lecture Presentations Multimodal Dataset: Towards Understanding Multimodality in Educational VideosDong Won Lee, Chaitanya Ahuja, Paul Pu Liang, Sanika Natu et al.ICCV 2023 · 20 citations
- VisualWebInstruct: Scaling up Multimodal Instruction Data through Web SearchYiming Jia, Jiachen Li, Xiang Yue, Bo Li et al.EMNLP 2025 · 29 citations
- MemeCap: A Dataset for Captioning and Interpreting MemesEunjeong Hwang, Vered ShwartzEMNLP 2023 · 14 citations
