Skip to contents

The aim of the package is to provide an implementation of the G-means algorithm in R. The G-means algorithm is a clustering algorithm that extends the k-means algorithm by automatically determining the number of clusters. The algorithm was introduced by Hamerly and Elkan (2003).

Installation

You can install the development version of gmeans from GitHub with:

# install.packages("pak")
pak::pak("m-muecke/gmeans")

Usage

The Old Faithful geyser has short and long eruptions, and gmeans() finds these two groups without being told how many to look for:

library(gmeans)

km <- gmeans(faithful)
summary(km)
#> G-means clustering with 2 clusters (k_init = 1, k_max = 10, level = 0.0001)
#> 
#>  cluster size withinss eruptions waiting
#>        1  172     5446     4.298   80.28
#>        2  100     3456     2.094   54.75
#> 
#> Total SS: 50440, within SS: 8902, between SS: 41538 (82.35% of total)

When to use gmeans

Use gmeans() when the number of clusters is unknown and the clusters are roughly Gaussian. The algorithm splits a cluster only when an Anderson-Darling test rejects normality, so the number of clusters follows from the data and a single significance level rather than from a grid search over k.

The default significance level of 0.0001 follows Hamerly and Elkan (2003) and keeps Gaussian clusters from being split. With small clusters of fewer than about 50 points the test has little power to detect a split, so a larger level such as 0.01 can work better.

  • mlr3cluster: cluster analysis for the mlr3 ecosystem, which provides G-means as the clust.gmeans learner.
  • nortest: R package for testing the composite hypothesis of normality.