Comparison of K-Means and Agglomerative Clustering on Mall Customers Dataset
Experiment
Comparison of K-Means and Agglomerative Clustering on Mall Customers Dataset
๐ฏ Objective
To apply and compare:
- K-Means (Partitional Clustering)
- Agglomerative (Hierarchical Clustering)
using evaluation metrics and visualization.
๐ Theory
๐น Partitional Clustering (K-Means)
- Divides data into K clusters
- Minimizes within-cluster variance
๐น Hierarchical Clustering (Agglomerative)
- Builds clusters bottom-up
- Produces a tree structure (dendrogram)
๐น Evaluation Metrics
1. Inertia (WCSS)
- Used for K-Means only
2. Silhouette Score
- Used for both methods
- Higher value → better clustering
๐งพ Dataset: Mall Customers
Typical features:
- CustomerID
- Gender
- Age
- Annual Income (k$)
- Spending Score (1–100)
๐ We use:
- Annual Income
- Spending Score
๐ป Program (Python Code)
๐ Output
K-Means Inertia: 65.56840815571681
K-Means Silhouette: 0.5546571631111091
Agglomerative Silhouette: 0.5538089226688662
๐ Observations
๐น K-Means
- Produces compact, spherical clusters
- Faster and scalable
๐น Agglomerative
- Captures hierarchical relationships
- More flexible cluster shapes
| Method | Inertia | Silhouette |
|---|---|---|
| K-Means | Low | ~0.55 |
| Agglomerative | — | ~0.55 |
๐ Result
Both clustering algorithms were applied successfully.
K-Means showed slightly better compactness, while Agglomerative provided hierarchical insights.
⚖️ Comparison
| Feature | K-Means | Agglomerative |
|---|---|---|
| Type | Partitional | Hierarchical |
| Speed | Fast | Slow |
| Scalability | High | Low |
| Shape handling | Spherical | Flexible |
| Output | Flat clusters | Tree (dendrogram) |
-
K-Means
- Efficient for large datasets
- Requires predefined K
-
Agglomerative
- More interpretable (dendrogram)
- Computationally expensive
๐ Choice depends on:
- Dataset size
- Cluster shape
- Need for hierarchy


Comments
Post a Comment