KNN Classification using Real Breast Cancer Dataset with Performance Evaluation
Experiment
Title
KNN Classification using Real Dataset with Performance Evaluation
๐ฏ Objective
- To implement K-Nearest Neighbors (KNN) using a real dataset
-
To evaluate model using:
- Precision
- Recall
- F1-score
- Accuracy
๐ Dataset
We use the Breast Cancer Dataset from sklearn:
-
Binary classification:
- 0 → Malignant
- 1 → Benign
- Features: Tumor characteristics (radius, texture, etc.)
๐ Theory
๐น KNN
- Classifies based on nearest neighbors
-
Uses distance metrics like:
- Euclidean distance
๐น Evaluation Metrics
- Precision → correctness of positive predictions
- Recall → ability to find all positives
- F1-score → balance of precision & recall
๐ป Complete Python Program
๐ Output
๐ Observations
- High accuracy (~90%+)
-
Good balance of:
- Precision
- Recall
- Class 1 (benign) often easier to classify
๐ Confusion Matrix Interpretation
| Predicted 0 | Predicted 1 | |
|---|---|---|
| Actual 0 | True Negatives | False Positives |
| Actual 1 | False Negatives | True Positives |
๐งช Lab Tasks
Task 1
Change K:
Task 2
Normalize data:
Task 3
Compare performance before and after scaling
๐ Key Insights
| Factor | Effect |
|---|---|
| Small K | Overfitting |
| Large K | Underfitting |
| Scaling | Improves KNN performance |
Result
- KNN performs well on real datasets
-
Sensitive to:
- Feature scaling
- Choice of K
- Evaluation metrics give better insight than accuracy alone
Comments
Post a Comment