Lab Assignment -11

 

Assignment-11- week 1 October

Learning Objective: Learn Clustering Techniques ( Unsupervised Learning)

1.K-Means Clustering

Problem Statement

A shopping mall wants to group its customers into different customer segments based on their Annual Income and Spending Score.

The following data is collected from 12 customers.

CustomerAnnual Income (₹ Lakhs)Spending Score
C12.02.0
C22.52.5
C33.02.0
C43.53.0
C56.07.0
C66.56.5
C77.07.5
C87.57.0
C99.02.0
C109.52.5
C1110.02.0
C1210.53.0

Tasks

  1. Represent the above data using a NumPy array or Pandas DataFrame.
  2. Plot the customers on a 2-D scatter plot, with:
    • Annual Income on the X-axis
    • Spending Score on the Y-axis.
  3. Apply the K-Means clustering algorithm with K = 3.
  4. Initialize three centroids and perform the following steps manually for at least the first two iterations:
    • Calculate the Euclidean distance between each point and each centroid.
    • Assign each point to its nearest centroid.
    • Recalculate the centroid of each cluster.
    • Repeat until the clusters become stable.
  5. Implement the same K-Means clustering using KMeans from Scikit-learn.
  6. Display the final:
    • Cluster assignment of every customer
    • Cluster centroids
  7. Plot the final clusters and their centroids using different markers/colors.
  8. Compare the manually obtained clusters with the Scikit-learn result.

Additional Tasks
  1. Repeat the experiment for K = 2 and K = 4.
  2. Compare the resulting clusters.
  3. Calculate the inertia (within-cluster sum of squares) for K = 2, 3, 4 and 5.
  4. Plot the Elbow Curve and determine a suitable value of K.
  5. Find the suitable value for K using silhouette method.

2.Finding the Optimal Number of Clusters using Elbow and Silhouette Methods

Problem Statement

Generate a random two-dimensional dataset containing 100 data points and use the K-Means clustering algorithm to determine the optimal number of clusters.

The optimal value of K should be determined using both:

  1. Elbow Method
  2. Silhouette Analysis

Tasks

  1. Generate 100 random 2-D data points using NumPy. You may use:
np.random.seed(42)
X = np.random.rand(100, 2) * 10
  1. Visualize the generated data using a scatter plot.
  2. Apply K-Means clustering for values of K from 2 to 10.
  3. For each value of K, calculate the inertia (Within-Cluster Sum of Squares).
  4. Plot:

K vs. Inertia

and identify the approximate elbow point.

  1. Calculate the Silhouette Score for each value of K from 2 to 10.
  2. Plot:

K vs. Silhouette Score

  1. Identify the value of K having the highest Silhouette Score.
  2. Compare the K obtained from:
    • Elbow Method
    • Silhouette Method
  3. Select a suitable value of K based on the two methods and apply K-Means clustering using the selected K.
  4. Visualize the final clusters and their centroids.

3.Experiment: Agglomerative Hierarchical Clustering and Dendrogram

Problem Statement

A shopping mall wants to group its customers based on their Annual Income and Spending Score. Unlike K-Means, the company does not know the number of clusters in advance.

Use Agglomerative Hierarchical Clustering to identify natural groups among the customers and visualize the clustering process using a dendrogram.

Sample Dataset

CustomerAnnual Income (₹ Lakhs)Spending Score
C12.02.0
C22.52.5
C33.02.0
C43.53.0
C56.07.0
C66.56.5
C77.07.5
C87.57.0
C99.02.0
C109.52.5
C1110.02.0
C1210.53.0

Tasks

  1. Create the dataset using a Pandas DataFrame.
  2. Plot the customers using a scatter plot.
  3. Calculate the Euclidean distance between the data points.
  4. Perform Agglomerative Hierarchical Clustering using:
    • n_clusters = 3
    • single linkage
  5. Display the cluster assigned to each customer.
  6. Visualize the resulting clusters using a scatter plot.
  7. Construct a dendrogram .
  8. Label the dendrogram with the customer IDs.
  9. Draw a horizontal line on the dendrogram to identify an appropriate number of clusters.
  10. Compare the clusters obtained from the dendrogram with those obtained using AgglomerativeClustering from Scikit-learn.

Additional Task

Repeat the clustering using different linkage methods:

  • Single linkage
  • Complete linkage
  • Average linkage
  • Ward linkage

Compare the resulting dendrograms and discuss how the choice of linkage method affects the clusters.

4.Implement and apply K-means clustering to the Digits dataset. Experiment with different numbers of clusters and evaluate the clustering results using metrics such as inertia and silhouette score. Analyze how the choice of K affects clustering performance.

Tasks:

● Load and preprocess the Digits dataset.
● Implement K-means clustering with various numbers of clusters.
● Evaluate clustering performance using inertia and silhouette score.
● Analyze the impact of the number of clusters on clustering quality.

5.Implement and compare hierarchical (agglomerative) and partitional (K-means) clustering algorithms on the Mall Customers dataset. Discuss the strengths and weaknesses of each method based on clustering results and evaluation metrics.

Tasks:

● Load and preprocess the Mall Customers dataset.
● Apply both hierarchical (agglomerative) and K-means clustering.
● Compare results using metrics such as inertia, silhouette score, and clustering visualization.
● Discuss the advantages and disadvantages of each clustering method.

Comments

Popular posts from this blog

Machine Learning Lab PCCSL508 Semester 5 KTU CS 2024 Scheme manual - Dr Binu V P

Lab Assignment-2

Lab Assignment-1