It's About Time: Analog Clock Reading in the Wild
Charig Yang, Weidi Xie, Andrew Zisserman
摘要
In this paper, we present a framework for reading analog clocks in natural images or videos. Specifically, we make the following contributions: First, we create a scalable pipeline for generating synthetic clocks, significantly reducing the requirements for the labour-intensive annotations; Second, we introduce a clock recognition architecture based on spatial transformer networks (STN), which is trained end-to-end for clock alignment and recognition. We show that the model trained on the proposed synthetic dataset generalises towards real clocks with good accuracy, advocating a Sim2Real training regime; Third, to further reduce the gap between simulation and real data, we leverage the special property of "time", i.e. uniformity, to generate reliable pseudo-labels on real unlabelled clock videos, and show that training on these videos offers further improvements while still requiring zero manual annotations. Lastly, we introduce three benchmark datasets based on COCO, Open Images, and The Clock movie, with full annotations for time, accurate to the minute.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Generative Multimodal Models are In-Context LearnersQuan Sun, Yufeng Cui, Xiaosong Zhang, Fan Zhang 等CVPR 2024 · 被引用 71 次
- Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBenchFenfen Lin, Yesheng Liu, Haiyu Xu, Chen Yue 等CVPR 2026 · 被引用 7 次
- HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal ModelTao Wang, Changxu Cheng, Lingfeng Wang, Senda Chen 等ICCV 2025 · 被引用 5 次
- Tiny Scales, Great Challenges: The Limits of Multimodal LLMs in Scale RecognitionJihang Jin, Ronghao Chen, Hao Zhang, Ziyan Liu 等ACL 2026
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language ModelsMatt Deitke, Christopher Clark, Sangho Lee, Rohun Tripathi 等CVPR 2025
它引用的顶会 Paper2
相关 Paper
- SPTS: Single-Point Text SpottingDezhi Peng, Xinyu Wang, Yuliang Liu, Jiaxin Zhang 等ACM MM 2022 · 被引用 65 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- MatchTime: Towards Automatic Soccer Game Commentary GenerationJiayuan Rao, Haoning Wu, Chang Liu, Yanfeng Wang 等EMNLP 2024 · 被引用 14 次
- WALT3D: Generating Realistic Training Data from Time-Lapse Imagery for Reconstructing Dynamic Objects Under OcclusionKhiem Vuong, N. Dinesh Reddy, Robert Tamburo, Srinivasa G. NarasimhanCVPR 2024 · 被引用 1 次
- Towards Omni-Supervised Face Alignment for Large Scale Unlabeled VideosCongcong Zhu, Hao Liu, Zhenhua Yu, Xuehong SunAAAI 2020 · 被引用 11 次
