Video Face Clustering With Unknown Number of Clusters
Makarand Tapaswi, Marc T. Law, Sanja Fidler
Abstract
Understanding videos such as TV series and movies requires analyzing who the characters are and what they are doing. We address the challenging problem of clustering face tracks based on their identity. Different from previous work in this area, we choose to operate in a realistic and difficult setting where: (i) the number of characters is not known a priori; and (ii) face tracks belonging to minor or background characters are not discarded. To this end, we propose Ball Cluster Learning (BCL), a supervised approach to carve the embedding space into balls of equal size, one for each cluster. The learned ball radius is easily translated to a stopping criterion for iterative merging algorithms. This gives BCL the ability to estimate the number of clusters as well as their assignment, achieving promising results on commonly used datasets. We also present a thorough discussion of how existing metric learning literature can be adapted for this task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 495404f7-1cfa-4042-9626-db4f1ee2597dCited by top-tier papers13
- Deep Open Intent Classification with Adaptive Decision BoundaryHanlei Zhang, Hua Xu, Ting-En LinAAAI 2021 · 127 citations
- A Theoretical Analysis of the Number of Shots in Few-Shot LearningTianshi Cao, Marc T. Law, Sanja FidlerICLR 2020 · 75 citations
- DeepDPM: Deep Clustering With an Unknown Number of ClustersMeitar Ronen, Shahaf E. Finder, Oren FreifeldCVPR 2022 · 66 citations
- AutoAD II: The Sequel - Who, When, and What in Movie Audio DescriptionTengda Han, Max Bain, Arsha Nagrani, Gül Varol et al.ICCV 2023 · 55 citations
- AVA-AVD: Audio-visual Speaker Diarization in the WildEric Zhongcong Xu, Zeyang Song, Satoshi Tsutsui, Chao Feng et al.ACM MM 2022 · 34 citations
Related papers
- Robust Actor Recognition in Entertainment Multimedia at ScaleAbhinav Aggarwal, Yash Pandya, Lokesh A. Ravindranathan, Laxmi S. Ahire et al.ACM MM 2022 · 4 citations
- Learned Trajectory Embedding for Subspace ClusteringYaroslava Lochman, Carl Olsson, Christopher ZachCVPR 2024 · 5 citations
- Tracklet Self-Supervised Learning for Unsupervised Person Re-IdentificationGuile Wu, Xiatian Zhu, Shaogang GongAAAI 2020 · 97 citations
- CLIP-Cluster: CLIP-Guided Attribute Hallucination for Face ClusteringShuai Shen, Wanhua Li, Xiaobing Wang, Dafeng Zhang et al.ICCV 2023 · 19 citations
- VideoCutLER: Surprisingly Simple Unsupervised Video Instance SegmentationXudong Wang, Ishan Misra, Ziyun Zeng, Rohit Girdhar et al.CVPR 2024
