arxivcs.LGcs.DS2026-07-06
Sensitivity Sampling with Predictions for k-Means Clustering
Cristian Boldrin, Fabio Vandin
We study the problem of k-means clustering on large datasets. The state-of-the-art for the problem is given by coresets-based approaches, which build small weighted summaries of the input and derive approximate solutions with rigorous quality guarantees from them. One of the most…