Counterintuitive Behavior of Clustering Quality

Findings for K-Means on Synthetic and Real Data

Conference Paper (2025)
Author(s)

Marco Loog (Radboud Universiteit Nijmegen)

Jesse H. Krijthe (TU Delft - Electrical Engineering, Mathematics and Computer Science)

Manuele Bicego (Università degli Studi di Verona)

Research Group
Pattern Recognition and Bioinformatics
DOI related publication
https://doi.org/10.1007/978-3-031-91398-3_12 Final published version
More Info
expand_more
Publication Year
2025
Language
English
Research Group
Pattern Recognition and Bioinformatics
Pages (from-to)
154-166
Publisher
Springer
ISBN (print)
9783031913976
Event
23rd International Symposium on Intelligent Data Analysis, IDA 2025 (2025-05-07 - 2025-05-09), Konstanz, Germany
Downloads counter
34
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

Little is known about how the quality of a clustering changes when changing the size of the set used to determine the clustering model. We show that, for K-means clustering, the relationship between dataset size and clustering quality can display counterintuitive behavior. Notably, the quality can significantly deteriorate with more data to build the model. More generally, using artificial datasets and data from bioinformatics, we uncover a variety of learning curve behaviors for K-means. Our results clearly illustrate that the training sample size can have a nontrivial influence on the clustering performance. Our findings should appeal to both the clustering practitioner and the clustering researcher concerned with developing basic insights.

Files

978-3-031-91398-3_12.pdf
(pdf | 1.25 Mb)
- Embargo expired in 02-11-2025
– Personal use only – Dutch Copyright Act (Article 25fa)