Learning Precise Temporal Point Event Detection with Misaligned Labels
Julien Schroeter, Kirill A. Sidorov, A. David Marshall
Abstract
This work addresses the problem of robustly learning precise temporal point event detection despite only having access to poorly aligned labels for training. While standard (cross entropy-based) methods work well in noise-free setting, they often fail when labels are unreliable since they attempt to strictly fit the annotations. A common solution to this drawback is to transform the point prediction problem into a distribution prediction problem. However, we show that this approach raises several issues that negatively affect the robust learning of temporal localization. Thus, in an attempt to overcome these shortcomings, we introduce a simple and versatile training paradigm combining soft localization learning with counting-based sparsity regularization. In fact, unlike its counterparts, our approach allows to directly infer clear-cut point predictions in an end-to-end fashion while relaxing the reliance of the training on the exact position of labels. We achieve state-of-the-art performance against standard benchmarks in a number of challenging experiments (e.g., detection of instantaneous events in videos and music transcription) by simply replacing the original loss function with our novel alternative---without any additional fine-tuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Related papers
- Learnable Negative Proposals Using Dual-Signed Cross-Entropy Loss for Weakly Supervised Video Moment LocalizationSunoh Kim, Daeho Um, Hyunjun Choi, Jin Young ChoiACM MM 2024 · 5 citations
- You Can even Annotate Text with Voice: Transcription-only-Supervised Text SpottingJingqun Tang, Su Qiao, Benlei Cui, Yuhang Ma et al.ACM MM 2022 · 22 citations
- PointTAD: Multi-Label Temporal Action Detection with Learnable Query PointsJing Tan, Xiaotong Zhao, Xintian Shi, Bin Kang et al.NeurIPS 2022 · 41 citations
- CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event LocalizationJinxing Zhou, Ziheng Zhou, Yanghao Zhou, Yuxin Mao et al.AAAI 2026 · 4 citations
- Boosting Point-Supervised Temporal Action Localization through Integrating Query Reformation and Optimal TransportMengnan Liu, Le Wang, Sanping Zhou, Kun Xia et al.CVPR 2025
