iTagPDF: Towards Finally Automating PDF Accessibility
Peya Mowar, Aaron Steinfeld, Jeffrey P. Bigham
Abstract
Most academic research is ultimately disseminated through documents in the PDF format. This format has advantages in flexibility and portability, but presents challenges for accessibility that have stubbornly resisted solutions despite decades of attempts. Tagging PDFs is hard to automate because tags are currently generated visually, not semantically, which makes the output cluttered and manual correction tedious and error-prone. Ironically, this semantic structure already exists during authoring but is discarded during PDF rendering. This raises an obvious question, can we use this lost semantic information to better automate tagging in PDFs? In this paper, we develop iTagPDF, a system that refines generated metadata using the semantics in the source documents of research papers. We demonstrate that the metadata generated by iTagPDF already surpasses what authors currently submit to ACM conferences on many criteria. Our approach represents a concrete step toward finally automating accessibility remediation in research paper PDFs.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Towards More Accessible Scientific PDFs for People with Visual Impairments: Step-by-Step PDF Remediation to Improve Tag AccuracyFelix Maximilian Schmitt-Koopmann, Elaine May Huang, Hans-Peter Hutter, Alireza DarvishyCHI 2025 · 6 citations
- TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX ReconstructionChengye Wang, Lin Fu, Zexi Kuang, Yilun ZhaoACL 2026
- PDF Mirage: Content Masking Attack Against Information-Based Online ServicesIan D. Markwood, Dakun Shen, Yao Liu, Zhuo LuUSENIX Security 2017 · 31 citations
- LATTE: Improving Latex Recognition for Tables and Formulae with Iterative RefinementNan Jiang, Shanchao Liang, Chengxiao Wang, Jiannan Wang et al.AAAI 2025 · 8 citations
- Screen Parsing: Towards Reverse Engineering of UI Models from ScreenshotsJason Wu, Xiaoyi Zhang, Jeffrey Nichols, Jeffrey P. BighamUIST 2021 · 62 citations
