K-means takes a list of numbers and sorts them into k groups. The groups are not something you give it, the algorithm works them out on its own. That makes it an unsupervised learning task, nothing in the data tells the algorithm what the answer is.

You can cluster it automatically with the kmeans algorithm.

In the kmeans algorithm, k is the number of clusters.

Clustering is an _unsupervised machine learning task. _ Everything is automatic.

kmeans data

We always start with data. This is our observed data, simply a list of values.
We plot all of the observed data in a scatter plot.

 # clustering dataset
from sklearn.cluster import KMeans
from sklearn import metrics
import numpy as np
import matplotlib.pyplot as plt

x1 = np.array([3, 1, 1, 2, 1, 6, 6, 6, 5, 6, 7, 8, 9, 8, 9, 9, 8])
x2 = np.array([5, 4, 6, 6, 5, 8, 6, 7, 6, 7, 1, 2, 1, 2, 3, 2, 3])

plt.plot()
plt.xlim([0, 10])
plt.ylim([0, 10])
plt.title('Dataset')
plt.scatter(x1, x2)
plt.show()

Result:
kmeans dataset

kmeans clustering example

We will cluster the observations automatically.

K can be determined using the elbow method, but in this example we’ll set K ourselves.

Note: K is always a positive integer. We cannot have -1 clusters (k).

k-means puts every observation in the group with the nearest centroid. Then it moves each centroid to the middle of its group and does it again.

Each observation belong to the cluster with the nearest mean.

 # clustering dataset
from sklearn.cluster import KMeans
from sklearn import metrics
import numpy as np
import matplotlib.pyplot as plt

x1 = np.array([3, 1, 1, 2, 1, 6, 6, 6, 5, 6, 7, 8, 9, 8, 9, 9, 8])
x2 = np.array([5, 4, 6, 6, 5, 8, 6, 7, 6, 7, 1, 2, 1, 2, 3, 2, 3])

plt.plot()
plt.xlim([0, 10])
plt.ylim([0, 10])
plt.title('Dataset')
plt.scatter(x1, x2)
plt.show()

# create new plot and data
plt.plot()
X = np.array(list(zip(x1, x2))).reshape(len(x1), 2)
colors = ['b', 'g', 'r']
markers = ['o', 'v', 's']

# KMeans algorithm
K = 3
kmeans_model = KMeans(n_clusters=K).fit(X)

plt.plot()
for i, l in enumerate(kmeans_model.labels_):
plt.plot(x1[i], x2[i], color=colors[l], marker=markers[l],ls='None')
plt.xlim([0, 10])
plt.ylim([0, 10])

plt.show()

Result:
kmeans clustering algorithm

If you see the above result, Kmeans has clustered the observations automatically.

What happens inside

  • Pick k starting points at random. These are the first centroids.
  • Every observation goes to the centroid closest to it.
  • Move each centroid to the average of the points assigned to it.
  • Repeat the last two steps until the centroids stop moving.

In scikit-learn that is KMeans(n_clusters=k) followed by fit() and predict(). The result is in kmeans.labels_ and the centres are in kmeans.cluster_centers_.

K-means assumes clusters that are roughly round and about the same size. Points far from the others pull their centroid away, and a very different spread in one cluster does not get a proper centroid of its own. You have to set k yourself, the elbow method helps.

Want more practice on this? Practise this on PyChallenge, there are short browser exercises you can run right after reading.