Implementation of K-Means Clustering Using Scikit-Learn
Experiment
Implementation of K-Means Clustering Using Scikit-Learn
๐ฏ Objective
To implement the K-Means clustering algorithm using scikit-learn and visualize how data points are grouped into clusters.
๐ Theory
๐น Clustering
Clustering is an unsupervised learning technique used to group similar data points based on their features.
๐น K-Means in Scikit-Learn
The KMeans class from sklearn.cluster provides an efficient implementation of the K-Means algorithm.
๐น Working Principle
- Specify number of clusters K
- Algorithm initializes centroids (random or k-means++)
- Assigns points to nearest centroid
- Updates centroids iteratively
- Stops when centroids stabilize
๐น Objective Function
K-Means minimizes:
๐งพ Sample Dataset
| Point | X | Y |
|---|---|---|
| P1 | 2 | 3 |
| P2 | 3 | 4 |
| P3 | 3 | 3 |
| P4 | 8 | 7 |
| P5 | 7 | 8 |
| P6 | 8 | 8 |
| P7 22 | 30 |
๐ป Program (Python Code Using Scikit-Learn)
๐ Output
๐น Cluster Labels
๐น Centroids
๐ Result
The K-Means algorithm using scikit-learn successfully grouped the dataset into 3 clusters, and the centroids were computed automatically.
scikit-learnsimplifies implementation of K-Means- Efficient for large datasets
- Automatically handles centroid initialization (k-means++)

Comments
Post a Comment