Lune

FOCS2023顶会

Deterministic Clustering in High Dimensional Spaces: Sketches and Approximation

Vincent Cohen-Addad, David Saulpic, Chris Schwiegelshohn

2023年份
3被引次数
6顶会引用

摘要

In all state-of-the-art sketching and coreset techniques for clustering, as well as in the best known fixed-parameter tractable approximation algorithms, randomness plays a key role. For the classic k-median and k-means problems, there are no known deterministic dimensionality reduction procedure or coreset construction that avoid an exponential dependency on the input dimension d, the precision parameter ε−1\varepsilon^{-1} or k. Furthermore, there is no coreset construction that succeeds with probability 1−1/n1-1/n and whose size does not depend on the number of input points, n. This has led researchers in the area to ask what is the power of randomness for clustering sketches [Feldman WIREs Data Mining Knowl. Discov’20].Similarly, the best approximation ratio achievable deterministically without a complexity exponential in the dimension are 1+21+\sqrt{2} for k-median [Cohen-Addad, Esfandiari, Mirrokni, Narayanan, STOC’22] and 6.12903 for k-means [Grandoni, Ostrovsky, Rabani, Schulman, Venkat, Inf. Process. Lett.’22]. Those are the best results, even when allowing a complexity FPT in the number of clusters k: this stands in sharp contrast with the (1+ε)(1+\varepsilon)-approximation achievable in that case, when allowing randomization.In this paper, we provide deterministic sketches constructions for clustering, whose size bounds are close to the best-known randomized ones. We show how to compute a dimension reduction onto ε−O(1)log⁡k\varepsilon^{-O(1)} \log k dimensions in time kO(ε−O(1)+log⁡log⁡k)k^{O\left(\varepsilon^{-O(1)}+\log \log k\right)} poly (nd)(n d), and how to build a coreset of size O(k2log⁡3kε−O(1))O\left(k^{2} \log ^{3} k \varepsilon^{-O(1)}\right) in time 2εO(1)klog⁡3k+kO(ε−O(1)+log⁡log⁡k)2^{\varepsilon^{O(1)} k \log ^{3} k}+k^{O\left(\varepsilon^{-O(1)}+\log \log k\right)} poly (nd)(n d). In the case where k is small, this answers an open question of [Feldman WIDM’20] and [Munteanu and Schwiegelshohn, Künstliche Intell. ’18] on whether it is possible to efficiently compute coresets deterministically.We also construct a deterministic algorithm for computing (1+(1+ ε)\varepsilon)-approximation to k-median and k-means in high dimensional Euclidean spaces in time 2k2log⁡3k/εO(1)2^{k^{2} \log ^{3} k / \varepsilon^{O(1)}} poly (nd)(n d), close to the best randomized complexity of 2(k/ε)O(1)2^{(k / \varepsilon)^{O(1)}} nd (see [Kumar, Sabharwal, Sen, JACM 10] and [Bhattacharya, Jaiswal, Kumar, TCS’18]).Furthermore, our new insights on sketches also yield a randomized coreset construction that uses uniform sampling, that immediately improves over the recent results of [Braverman et al. FOCS ’22] by a factor k.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper6

问问它们各自怎么用它

它引用的顶会 Paper12

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖