Reading Recognition in the Wild
Charig Yang, Samiul Alam, Shakhrul Iman Siam, Michael J. Proulx, Lambert Mathias, Kiran K. Somasundaram, Luis Pesqueira, James Fort, Sheroze Sheriffdeen, Omkar M. Parkhi, Carl Yuheng Ren, Mi Zhang
Abstract
To enable egocentric contextual AI in always-on smart glasses, it is crucial to be able to keep a record of the user's interactions with the world, including during reading. In this paper, we introduce a new task of reading recognition to determine when the user is reading. We first introduce the first-of-its-kind large-scale multimodal Reading in the Wild dataset, containing 100 hours of reading and non-reading videos in diverse and realistic scenarios. We then identify three modalities (egocentric RGB, eye gaze, head pose) that can be used to solve the task, and present a flexible transformer model that performs the task using these modalities, either individually or combined. We show that these modalities are relevant and complementary to the task, and investigate how to efficiently and effectively encode each modality. Additionally, we show the usefulness of this dataset towards classifying types of reading, extending current reading understanding studies conducted in constrained settings to larger scale, diversity and realism. Code, model, and data will be public. Figure 1: Am I reading? The left figure shows a timeline as the user navigates the world. We aim to solve the task of reading recognition to enable AI assistants in always-on wearables. We identify three modalities: eye gaze (in colored dot patterns), RGB crop around gaze (in red box), and inertial sensors performs the task to high accuracy (with Prediction and GT shown). Images from our Reading in the Wild dataset, which features 100 hours of diverse reading and non-reading activities in real-world settings, with examples shown in the right.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c6a58eb-ae75-46e1-8a17-24b3231b84f7Cited by top-tier papers1
Ask how each one uses itBuilds on4
- EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action RecognitionEvangelos Kazakos, Arsha Nagrani, Andrew Zisserman, Dima DamenICCV 2019 · 395 citations
- GazePrompt: Enhancing Low Vision People's Reading Experience with Gaze-Aware AugmentationsRu Wang, Zach Potter, Yun Ho, Daniel Killough et al.CHI 2024 · 13 citations
- Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person PerspectivesKristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani et al.CVPR 2024
- Learning Video Representations from Large Language ModelsYue Zhao, Ishan Misra, Philipp Krähenbühl, Rohit GirdharCVPR 2023
Related papers
- OSMO: Open-vocabulary Self-eMOtion TrackingMohamed Abdelfattah, Bugra Tekin, Fadime Sener, Necati Cihan Camgoz et al.CVPR 2026
- CASES: A Cognition-Aware Smart Eyewear System for Understanding How People ReadXiangyao Qi, Qi Lu, Wentao Pan, Yingying Zhao et al.UbiComp 2023 · 9 citations
- Multimodal Daily-Life Logging in Free-living Environment Using Non-Visual Egocentric Sensors on a SmartphoneKe Sun, Chunyu Xia, Xinyu Zhang, Hao Chen et al.UbiComp 2024 · 15 citations
- Seeing Conversations: Communication Context Identification in Egocentric VideoTobias Dorszewski, Jens HjortkjærCVPR 2026
- WearVox: An Egocentric Multichannel Voice Assistant Benchmark for WearablesZhaojiang Lin, Yong Xu, Kai Sun, Jing Zheng et al.ICLR 2026 · 11 citations
