SillyPutty: Improved clustering by optimizing the silhouette width.

Bombina, Polina; Tally, Dwayne; Abrams, Zachary B; Coombes, Kevin R

Bombina, Polina; Tally, Dwayne; Abrams, Zachary B; Coombes, Kevin R.

Afiliación

Bombina P; Department of Biostatistics, Data Science and Epidemiology, Georgia Cancer Center at Augusta University, Augusta, GA, United States of America.
Tally D; Department of Informatics, Indiana University, United States of America.
Abrams ZB; Division of Data Science and Biostatistics, Institute for Informatics, Washington University School of Medicine, Saint Louis, MO, United States of America.
Coombes KR; Department of Biostatistics, Data Science and Epidemiology, Georgia Cancer Center at Augusta University, Augusta, GA, United States of America.

PLoS One ; 19(6): e0300358, 2024.

Article en En | MEDLINE | ID: mdl-38848330

ABSTRACT

ABSTRACT

Clustering is an important task in biomedical science, and it is widely believed that different data sets are best clustered using different algorithms. When choosing between clustering algorithms on the same data set, reseachers typically rely on global measures of quality, such as the mean silhouette width, and overlook the fine details of clustering. However, the silhouette width actually computes scores that describe how well each individual element is clustered. Inspired by this observation, we developed a novel clustering method, called SillyPutty. Unlike existing methods, SillyPutty uses the silhouette width for individual elements as a tool to optimize the mean silhouette width. This shift in perspective allows for a more granular evaluation of clustering quality, potentially addressing limitations in current methodologies. To test the SillyPutty algorithm, we first simulated a series of data sets using the Umpire R package and then used real-workd data from The Cancer Genome Atlas. Using these data sets, we compared SillyPutty to several existing algorithms using multiple metrics (Silhouette Width, Adjusted Rand Index, Entropy, Normalized Within-group Sum of Square errors, and Perfect Classification Count). Our findings revealed that SillyPutty is a valid standalone clustering method, comparable in accuracy to the best existing methods. We also found that the combination of hierarchical clustering followed by SillyPutty has the best overall performance in terms of both accuracy and speed.

Availability:

The SillyPutty R package can be downloaded from the Comprehensive R Archive Network (CRAN).

Asunto(s)

Algoritmos; Análisis por Conglomerados; Humanos; Neoplasias/patología; Programas Informáticos

Texto completo

Imprimir

XML

PubMed Links

Buscar en Google

Texto completo: 1 Base de datos: MEDLINE Asunto principal: Algoritmos Idioma: En Revista: PLoS ONE (Online) / PLoS One / PLos ONE Asunto de la revista: CIENCIA / MEDICINA Año: 2024 Tipo del documento: Article

Texto completo

Imprimir

XML

PubMed Links

Buscar en Google