Implementation of K-Means Clustering on Randomly Generated Data with 5 Clusters

 

Experiment

Implementation of K-Means Clustering on Randomly Generated Data with 5 Clusters


🎯 Objective

To generate sample data containing five clusters, apply the K-Means clustering algorithm, and visualize the clustering results.


📘 Theory

🔹 K-Means Clustering

K-Means is an unsupervised learning algorithm that groups data points into K clusters based on similarity.

The algorithm attempts to minimize the distance between data points and their assigned cluster centroids.


🔹 Working Steps of K-Means

  1. Select the number of clusters (K)
  2. Initialize K centroids
  3. Compute distance of each point from all centroids
  4. Assign points to nearest centroid
  5. Compute new centroids
  6. Repeat until convergence

🔹 Objective Function

K-Means minimizes:

J=∑i=1K∑x∈Ci∣∣x−μi∣∣2J=\sum_{i=1}^{K}\sum_{x\in C_i}||x-\mu_i||^2

Where:

  • KK = number of clusters
  • CiC_i = cluster ii
  • μi\mu_i= centroid of cluster ii

🧾 Dataset Description

Instead of using an external dataset, data points are generated randomly.

Characteristics:

  • Total samples = 500
  • Number of clusters = 5
  • Features = 2
  • Cluster centers generated automatically

💻 Program

import numpy as np
import matplotlib.pyplot as plt
from sklearn.datasets import make_blobs
from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score

# ------------------------------------
# Step 1: Generate random dataset
# ------------------------------------

X, y = make_blobs(
n_samples=500,
centers=5,
n_features=2,
random_state=42
)

# ------------------------------------
# Step 2: Visualize original data
# ------------------------------------

plt.figure()

plt.scatter(X[:,0], X[:,1])

plt.title("Randomly Generated Dataset")
plt.xlabel("Feature 1")
plt.ylabel("Feature 2")

plt.show()

# ------------------------------------
# Step 3: Apply K-Means
# ------------------------------------

k=5

kmeans = KMeans(
n_clusters=k,
init='k-means++',
random_state=42
)

labels = kmeans.fit_predict(X)

centroids = kmeans.cluster_centers_

# ------------------------------------
# Step 4: Compute metrics
# ------------------------------------

inertia = kmeans.inertia_

sil_score = silhouette_score(X,labels)

# ------------------------------------
# Step 5: Visualize clusters
# ------------------------------------

plt.figure()

plt.scatter(
X[:,0],
X[:,1],
c=labels
)

plt.scatter(
centroids[:,0],
centroids[:,1],
marker='X',
            color='red'
s=300
)

plt.title("K-Means Clustering with K=5")
plt.xlabel("Feature 1")
plt.ylabel("Feature 2")

plt.show()

# ------------------------------------
# Step 6: Display output
# ------------------------------------

print("Centroids:\n")
print(centroids)

print("\nInertia:",inertia)

print("\nSilhouette Score:",sil_score)

📊  Output




Centroids: [[-2.64387445 9.04103329] [-6.88093348 -6.9950069 ] [ 4.74346293 1.99434819] [-8.91925181 7.38961974] [ 2.13315564 4.35513167]] Inertia: 924.0989923406021 Silhouette Score: 0.678738720085253


📈 Result Analysis

Inertia

  • Measures compactness of clusters
  • Lower value indicates tighter clusters

Silhouette Score

Interpretation:

ScoreMeaning
Near +1    Well-separated clusters
Near 0    Overlapping clusters
Less than 0    Incorrect clustering

For this experiment:

Silhouette Score ≈ 0.7

This indicates good clustering quality.


📌 Result

K-Means clustering was successfully applied on randomly generated data consisting of five clusters. The algorithm identified clusters and computed centroids successfully.





Comments

Popular posts from this blog

Machine Learning Lab PCCSL508 Semester 5 KTU CS 2024 Scheme manual - Dr Binu V P

Lab Assignment-2

Lab Assignment-1