The aim of the package is to provide an implementation of the G-means algorithm in R. The G-means algorithm is a clustering algorithm that extends the k-means algorithm by automatically determining the number of clusters. The algorithm was introduced by Hamerly and Elkan (2003).
Installation
You can install the development version of gmeans from GitHub with:
# install.packages("pak")
pak::pak("m-muecke/gmeans")Usage
The Old Faithful geyser has short and long eruptions, and gmeans() finds these two groups without being told how many to look for:
library(gmeans)
km <- gmeans(faithful)
summary(km)
#> G-means clustering with 2 clusters (k_init = 1, k_max = 10, level = 0.0001)
#>
#> cluster size withinss eruptions waiting
#> 1 172 5446 4.298 80.28
#> 2 100 3456 2.094 54.75
#>
#> Total SS: 50440, within SS: 8902, between SS: 41538 (82.35% of total)When to use gmeans
Use gmeans() when the number of clusters is unknown and the clusters are roughly Gaussian. The algorithm splits a cluster only when an Anderson-Darling test rejects normality, so the number of clusters follows from the data and a single significance level rather than from a grid search over k.
The default significance level of 0.0001 follows Hamerly and Elkan (2003) and keeps Gaussian clusters from being split. With small clusters of fewer than about 50 points the test has little power to detect a split, so a larger level such as 0.01 can work better.
Related work
-
mlr3cluster: cluster analysis for the mlr3 ecosystem, which provides G-means as the
clust.gmeanslearner. - nortest: R package for testing the composite hypothesis of normality.