Implementation of K-Means Clustering Using Scikit-Learn

 

Experiment

Implementation of K-Means Clustering Using Scikit-Learn


๐ŸŽฏ Objective

To implement the K-Means clustering algorithm using scikit-learn and visualize how data points are grouped into clusters.


๐Ÿ“˜ Theory

๐Ÿ”น Clustering

Clustering is an unsupervised learning technique used to group similar data points based on their features.


๐Ÿ”น K-Means in Scikit-Learn

The KMeans class from sklearn.cluster provides an efficient implementation of the K-Means algorithm.


๐Ÿ”น Working Principle

  1. Specify number of clusters K
  2. Algorithm initializes centroids (random or k-means++)
  3. Assigns points to nearest centroid
  4. Updates centroids iteratively
  5. Stops when centroids stabilize

๐Ÿ”น Objective Function

K-Means minimizes:

J=∑i=1K∑x∈Ci∣∣x−ฮผi∣∣2J = \sum_{i=1}^{K} \sum_{x \in C_i} ||x - \mu_i||^2

๐Ÿงพ Sample Dataset

PointXY
P123
P234
P333
P487
P578
P688
P7               2230
P8               25     30

๐Ÿ’ป Program (Python Code Using Scikit-Learn)

import numpy as np import matplotlib.pyplot as plt from sklearn.cluster import KMeans # Step 1: Dataset X = np.array([ [2, 3], [3, 4], [3, 3], [8, 7], [7, 8], [8, 8],
[22,28], [25, 30] ]) # Step 2: Define number of clusters k = 3 # Step 3: Apply K-Means kmeans = KMeans(n_clusters=k, init='k-means++', random_state=0) kmeans.fit(X) # Step 4: Get labels and centroids labels = kmeans.labels_ centroids = kmeans.cluster_centers_ # Step 5: Visualization plt.scatter(X[:, 0], X[:, 1], c=labels) # Plot centroids plt.scatter(centroids[:, 0], centroids[:, 1], marker='X', s=200, c='red') plt.title("K-Means Clustering (Scikit-Learn)") plt.xlabel("X") plt.ylabel("Y") plt.show() # Step 6: Output print("Cluster Labels:", labels) print("Centroids:\n", centroids)

๐Ÿ“Š Output

๐Ÿ”น Cluster Labels 

Cluster Labels: [2 2 2 0 0 0 1 1]

๐Ÿ”น Centroids 

Centroids: [[ 7.66666667 7.66666667] [23.5 29. ] [ 2.66666667 3.33333333]]



๐Ÿ“Œ Result

The K-Means algorithm using scikit-learn successfully grouped the dataset into 3 clusters, and the centroids were computed automatically.

  • scikit-learn simplifies implementation of K-Means
  • Efficient for large datasets
  • Automatically handles centroid initialization (k-means++)

Comments

Popular posts from this blog

Machine Learning Lab PCCSL508 Semester 5 KTU CS 2024 Scheme manual - Dr Binu V P

Lab Assignment-2

Lab Assignment-1