K-Means Clustering on Digits Dataset with Performance Evaluation
Experiment
K-Means Clustering on Digits Dataset with Performance Evaluation
🎯 Objective
To apply K-Means clustering on the Digits dataset and evaluate clustering performance using:
- Inertia (WCSS)
- Silhouette Score
📘 Theory
🔹 Digits Dataset
- Contains 1797 samples of handwritten digits (0–9)
- Each image is 8×8 pixels (64 features)
- Used for classification and clustering tasks
🔹 K-Means Clustering
Partitions data into K clusters by minimizing within-cluster variance.
🔹 Evaluation Metrics
1. Inertia (WCSS)
- Lower value → better clustering
- Always decreases as K increases
2. Silhouette Score
- Range: -1 to 1
- Higher value → better clustering
💻 Program (Python Code)
📊 Expected Observations
🔹 Inertia Plot
- Decreases continuously as K increases
- Elbow typically around:
👉 K ≈ 8 to 10
🔹 Silhouette Plot
- Peaks at a certain K
- Often lower for very high K
👉 May not be exactly 10 (true classes)
📈 Analysis
🔹 Effect of K on Inertia
- K ↑ → Inertia ↓
- Reason: more clusters → tighter grouping
🔹 Effect of K on Silhouette Score
- Initially increases
- Then decreases after optimal K
🔹 Important Insight
👉 Even though dataset has 10 digits, optimal K may differ because:
- K-Means assumes spherical clusters
- Digits may overlap in feature space
📌 Result
- K-Means clustering was applied to the Digits dataset
-
Optimal K identified using:
- Elbow Method (Inertia)
- Silhouette Score
👉 Observed optimal K ≈ 8–10
- Inertia alone is not sufficient
- Silhouette provides better cluster validation
- Real-world datasets may not give perfect K


Comments
Post a Comment