Probabilistic Knowledge Distillation of Face Ensembles
Jianqing Xu, Shen Li, Ailin Deng, Miao Xiong, Jiaying Wu, Jiaxiang Wu, Shouhong Ding, Bryan Hooi
Abstract
Mean ensemble (i.e. averaging predictions from multiple models) is a commonly-used technique in machine learning that improves the performance of each individual model. We formalize it as feature alignment for ensemble in open-set face recognition and generalize it into Bayesian Ensemble Averaging (BEA) through the lens of probabilistic modeling. This generalization brings up two practical benefits that existing methods could not provide: (1) the uncertainty of a face image can be evaluated and further decomposed into aleatoric uncertainty and epistemic uncertainty, the latter of which can be used as a measure for out-of-distribution detection of faceness; (2) a BEA statistic provably reflects the aleatoric uncertainty of a face image, acting as a measure for face image quality to improve recognition performance. To inherit the uncertainty estimation capability from BEA without the loss of inference efficiency, we propose BEA-KD, a student model to distill knowledge from BEA. BEA-KD mimics the overall behavior of ensemble members and consistently outperforms SOTA knowledge distillation methods on various challenging benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- ID3: Identity-Preserving-yet-Diversified Diffusion Models for Synthetic Face RecognitionJianqing Xu, Shen Li, Jiaying Wu, Miao Xiong et al.NeurIPS 2024 · 33 citations
- Seeing Unseen: Discover Novel Biomedical Concepts via Geometry-Constrained Probabilistic ModelingJianan Fan, Dongnan Liu, Hang Chang, Heng Huang et al.CVPR 2024
- FreqDebias: Towards Generalizable Deepfake Detection via Consistency-Driven Frequency DebiasingHossein Kashiani, Niloufar Alipour Talemi, Fatemeh AfghahCVPR 2025
Builds on8
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- Probabilistic Face EmbeddingsYichun Shi, Anil K. JainICCV 2019 · 362 citations
- Sample and Computation Redistribution for Efficient Face DetectionJia Guo, Jiankang Deng, Alexandros Lattas, Stefanos ZafeiriouICLR 2022 · 173 citations
- Densely Guided Knowledge Distillation using Multiple Teacher AssistantsWonchul Son, Jaemin Na, Junyong Choi, Wonjun HwangICCV 2021 · 158 citations
- Agree to Disagree: Adaptive Ensemble Knowledge Distillation in Gradient SpaceShangchen Du, Shan You, Xiaojie Li, Jianlong Wu et al.NeurIPS 2020 · 144 citations
Related papers
- Face Alignment With Kernel Density Deep Neural NetworkLisha Chen, Hui Su, Qiang JiICCV 2019 · 34 citations
- From Risk to Uncertainty: Generating Predictive Uncertainty Measures via Bayesian EstimationNikita Kotelevskii, Vladimir Kondratyev, Martin Takác, Eric Moulines et al.ICLR 2025
- Provable Uncertainty Decomposition via Higher-Order CalibrationGustaf Ahdritz, Aravind Gollakota, Parikshit Gopalan, Charlotte Peale et al.ICLR 2025
- Cooperative Variance Estimation and Bayesian Neural Networks for Disentangling Aleatoric and Epistemic UncertaintiesJiaxiang Yi, Miguel BessaICML 2026 · 2 citations
- Bayesian Knowledge Distillation: A Bayesian Perspective of Distillation with Uncertainty QuantificationLuyang Fang, Yongkai Chen, Wenxuan Zhong, Ping MaICML 2024 · 10 citations
