DeFB: Decomposed Feature Learning for Real-Time Multi-Person Eyeblink Detection in Untrimmed In-the-Wild Videos
Jinfang Gan, Wenzheng Zeng, Yang Xiao, Xintao Zhang, Chaoyang Zheng, Ran Zhao, Ran Wang, Min Du, Zhiguo Cao
Abstract
Multi-person eyeblink detection in untrimmed in-the-wild videos is a recently emerged and challenging task. Due to its significant spatio-temporal fine-grained characteristics compared to general actions, we empirically find that general action detectors, though effective in general domains, struggle with this task (i.e., Blink-AP < 2%). Specialized eyeblink detection methods alleviate it through fine-grained spatio-temporal operations. SOTA method proposes a unified model combining instance-aware face localization and eyeblink detection through joint multi-task learning and feature sharing. While effective, it exhibits two critical limitations that may contribute to its unsatisfactory performance (i.e., Blink-AP=10.11%): (1) Face localization and eyeblink detection require distinct spatio-temporal feature granularities, making joint modeling in a unified feature space suboptimal. (2) Eyeblink task training could be largely affected by unstable face-eye feature learning under the joint training paradigm. To address this, we propose DeFB, a decomposed feature learning paradigm with favorable effectiveness and efficiency: (1) We model faces and eyes in granularity-specific feature spaces, which enhances fine-grained perception while reducing computational costs compared to a unified feature space. (2) To mitigate face-eye feature learning instability, we adopt an asynchronous learning mechanism where eye feature learning refines well-trained coarse face features, with shared queries acting as a bridge between stages to retain the efficient feature sharing of existing unified models. Compared with SOTA method, DeFB doubles the performance (Blink-AP: 24.65% v.s. 10.11%) while boosting efficiency by nearly 35%. DeFB can also be integrated as a plug-in to substantially augment the eyeblink detection capabilities of general action detectors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on13
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng et al.ICCV 2021 · 974 citations
- DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingFeng Li, Hao Zhang, Shilong Liu, Jian Guo et al.CVPR 2022 · 879 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- Fake it till you make it: face analysis in the wild using synthetic data aloneErroll Wood, Tadas Baltrusaitis, Charlie Hewitt, Sebastian Dziadzio et al.ICCV 2021 · 331 citations
Related papers
- Real-time Multi-person Eyeblink Detection in the Wild for Untrimmed VideoWenzheng Zeng, Yang Xiao, Sicheng Wei, Jinfang Gan et al.CVPR 2023
- MUP: Multi-granularity Unified Perception for Panoramic Activity RecognitionMeiqi Cao, Rui Yan, Xiangbo Shu, Jiachao Zhang et al.ACM MM 2023 · 12 citations
- GazeOnce: Real-Time Multi-Person Gaze EstimationMingfang Zhang, Yunfei Liu, Feng LuCVPR 2022 · 30 citations
- Real-time 3D neural facial animation from binocular videoChen Cao, Vasu Agrawal, Fernando De la Torre, Lele Chen et al.SIGGRAPH 2021 · 29 citations
- Temporal Action Localization with Cross Layer Task Decoupling and RefinementQiang Li, Di Liu, Jun Kong, Sen Li et al.AAAI 2025 · 3 citations
