KNN Classification using Real Breast Cancer Dataset with Performance Evaluation

 

Experiment

Title

KNN Classification using Real Dataset with Performance Evaluation


๐ŸŽฏ Objective

  • To implement K-Nearest Neighbors (KNN) using a real dataset
  • To evaluate model using:
    • Precision
    • Recall
    • F1-score
    • Accuracy

๐Ÿ“Š Dataset

We use the Breast Cancer Dataset from sklearn:

  • Binary classification:
    • 0 → Malignant
    • 1 → Benign
  • Features: Tumor characteristics (radius, texture, etc.)

๐Ÿ“š Theory

๐Ÿ”น KNN

  • Classifies based on nearest neighbors
  • Uses distance metrics like:
    • Euclidean distance

๐Ÿ”น Evaluation Metrics

  • Precision → correctness of positive predictions
  • Recall → ability to find all positives
  • F1-score → balance of precision & recall

๐Ÿ’ป Complete Python Program

# ------------------------------- # 1. Import Libraries # ------------------------------- from sklearn.datasets import load_breast_cancer from sklearn.model_selection import train_test_split from sklearn.neighbors import KNeighborsClassifier from sklearn.metrics import accuracy_score, classification_report, confusion_matrix # ------------------------------- # 2. Load Dataset # ------------------------------- data = load_breast_cancer() X = data.data y = data.target # ------------------------------- # 3. Train-Test Split # ------------------------------- X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42 ) # ------------------------------- # 4. Train KNN Model # ------------------------------- knn = KNeighborsClassifier(n_neighbors=5) knn.fit(X_train, y_train) # ------------------------------- # 5. Predictions # ------------------------------- y_pred = knn.predict(X_test) # ------------------------------- # 6. Evaluation # ------------------------------- print("Accuracy:", accuracy_score(y_test, y_pred)) print("\nClassification Report:") print(classification_report(y_test, y_pred)) print("\nConfusion Matrix:") print(confusion_matrix(y_test, y_pred))

๐Ÿ“ˆ  Output 

Accuracy: 0.956140350877193 Classification Report: precision recall f1-score support 0 1.00 0.88 0.94 43 1 0.93 1.00 0.97 71 accuracy 0.96 114 macro avg 0.97 0.94 0.95 114 weighted avg 0.96 0.96 0.96 114 Confusion Matrix: [[38 5] [ 0 71]]

๐Ÿ” Observations

  • High accuracy (~90%+)
  • Good balance of:
    • Precision
    • Recall
  • Class 1 (benign) often easier to classify

๐Ÿ“Š Confusion Matrix Interpretation

Predicted 0    Predicted 1
Actual 0    True Negatives    False Positives
Actual 1    False Negatives    True Positives

๐Ÿงช Lab Tasks

Task 1

Change K:

n_neighbors = 3, 7, 9

Task 2

Normalize data:

from sklearn.preprocessing import StandardScaler

Task 3

Compare performance before and after scaling


๐Ÿ“Š Key Insights

FactorEffect
Small K    Overfitting
Large K    Underfitting
Scaling    Improves KNN performance

Result

  • KNN performs well on real datasets
  • Sensitive to:
    • Feature scaling
    • Choice of K
  • Evaluation metrics give better insight than accuracy alone

Comments

Popular posts from this blog

Machine Learning Lab PCCSL508 Semester 5 KTU CS 2024 Scheme manual - Dr Binu V P

Lab Assignment-2

Lab Assignment-1