Lab Assignment -11
Assignment-11- week 1 October
Learning Objective: Learn Clustering Techniques ( Unsupervised Learning)
1.K-Means Clustering
Problem Statement
A shopping mall wants to group its customers into different customer segments based on their Annual Income and Spending Score.
The following data is collected from 12 customers.
| Customer | Annual Income (₹ Lakhs) | Spending Score |
|---|---|---|
| C1 | 2.0 | 2.0 |
| C2 | 2.5 | 2.5 |
| C3 | 3.0 | 2.0 |
| C4 | 3.5 | 3.0 |
| C5 | 6.0 | 7.0 |
| C6 | 6.5 | 6.5 |
| C7 | 7.0 | 7.5 |
| C8 | 7.5 | 7.0 |
| C9 | 9.0 | 2.0 |
| C10 | 9.5 | 2.5 |
| C11 | 10.0 | 2.0 |
| C12 | 10.5 | 3.0 |
Tasks
- Represent the above data using a NumPy array or Pandas DataFrame.
- Plot the customers on a 2-D scatter plot, with:
- Annual Income on the X-axis
- Spending Score on the Y-axis.
- Apply the K-Means clustering algorithm with K = 3.
- Initialize three centroids and perform the following steps manually for at least the first two iterations:
- Calculate the Euclidean distance between each point and each centroid.
- Assign each point to its nearest centroid.
- Recalculate the centroid of each cluster.
- Repeat until the clusters become stable.
- Implement the same K-Means clustering using
KMeansfrom Scikit-learn. - Display the final:
- Cluster assignment of every customer
- Cluster centroids
- Plot the final clusters and their centroids using different markers/colors.
- Compare the manually obtained clusters with the Scikit-learn result.
Additional Tasks
- Repeat the experiment for K = 2 and K = 4.
- Compare the resulting clusters.
- Calculate the inertia (within-cluster sum of squares) for K = 2, 3, 4 and 5.
- Plot the Elbow Curve and determine a suitable value of K.
- Find the suitable value for K using silhouette method.
2.Finding the Optimal Number of Clusters using Elbow and Silhouette Methods
Problem Statement
Generate a random two-dimensional dataset containing 100 data points and use the K-Means clustering algorithm to determine the optimal number of clusters.
The optimal value of K should be determined using both:
- Elbow Method
- Silhouette Analysis
Tasks
- Generate 100 random 2-D data points using NumPy. You may use:
np.random.seed(42) X = np.random.rand(100, 2) * 10
- Visualize the generated data using a scatter plot.
- Apply K-Means clustering for values of K from 2 to 10.
- For each value of K, calculate the inertia (Within-Cluster Sum of Squares).
- Plot:
K vs. Inertia
and identify the approximate elbow point.
- Calculate the Silhouette Score for each value of K from 2 to 10.
- Plot:
K vs. Silhouette Score
- Identify the value of K having the highest Silhouette Score.
-
Compare the K obtained from:
- Elbow Method
- Silhouette Method
- Select a suitable value of K based on the two methods and apply K-Means clustering using the selected K.
- Visualize the final clusters and their centroids.
Comments
Post a Comment