It's About Time: Analog Clock Reading in the Wild
Charig Yang, Weidi Xie, Andrew Zisserman
Abstract
In this paper, we present a framework for reading analog clocks in natural images or videos. Specifically, we make the following contributions: First, we create a scalable pipeline for generating synthetic clocks, significantly reducing the requirements for the labour-intensive annotations; Second, we introduce a clock recognition architecture based on spatial transformer networks (STN), which is trained end-to-end for clock alignment and recognition. We show that the model trained on the proposed synthetic dataset generalises towards real clocks with good accuracy, advocating a Sim2Real training regime; Third, to further reduce the gap between simulation and real data, we leverage the special property of "time", i.e. uniformity, to generate reliable pseudo-labels on real unlabelled clock videos, and show that training on these videos offers further improvements while still requiring zero manual annotations. Lastly, we introduce three benchmark datasets based on COCO, Open Images, and The Clock movie, with full annotations for time, accurate to the minute.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65889dfb-afb5-4059-859d-180cc78033eeCited by top-tier papers5
- Generative Multimodal Models are In-Context LearnersQuan Sun, Yufeng Cui, Xiaosong Zhang, Fan Zhang et al.CVPR 2024 · 71 citations
- Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBenchFenfen Lin, Yesheng Liu, Haiyu Xu, Chen Yue et al.CVPR 2026 · 7 citations
- HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal ModelTao Wang, Changxu Cheng, Lingfeng Wang, Senda Chen et al.ICCV 2025 · 5 citations
- Tiny Scales, Great Challenges: The Limits of Multimodal LLMs in Scale RecognitionJihang Jin, Ronghao Chen, Hao Zhang, Ziyan Liu et al.ACL 2026
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language ModelsMatt Deitke, Christopher Clark, Sangho Lee, Rohun Tripathi et al.CVPR 2025
Builds on2
Related papers
- SPTS: Single-Point Text SpottingDezhi Peng, Xinyu Wang, Yuliang Liu, Jiaxin Zhang et al.ACM MM 2022 · 65 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- MatchTime: Towards Automatic Soccer Game Commentary GenerationJiayuan Rao, Haoning Wu, Chang Liu, Yanfeng Wang et al.EMNLP 2024 · 14 citations
- WALT3D: Generating Realistic Training Data from Time-Lapse Imagery for Reconstructing Dynamic Objects Under OcclusionKhiem Vuong, N. Dinesh Reddy, Robert Tamburo, Srinivasa G. NarasimhanCVPR 2024 · 1 citation
- Towards Omni-Supervised Face Alignment for Large Scale Unlabeled VideosCongcong Zhu, Hao Liu, Zhenhua Yu, Xuehong SunAAAI 2020 · 11 citations
