Finding the Optimal Value of K in KNN using Grid Search

 

Experiment: 

Finding the Optimal Value of K in KNN using Grid Search

πŸ“Œ Aim

To implement K-Nearest Neighbors (KNN) Classification and determine the best value of K using Grid Search with Cross Validation.


🎯 Objectives

  • Understand KNN classification
  • Generate sample classification data
  • Implement KNN classifier
  • Use GridSearchCV to find optimal K
  • Evaluate classification performance
  • Visualize accuracy for different K values

πŸ“– Theory


πŸ”Ή K-Nearest Neighbors (KNN)

KNN is a supervised machine learning algorithm used for:

  • classification
  • regression

In KNN classification:

  • the class of a new point is determined by the majority class among its K nearest neighbors.

πŸ”Ή Working of KNN

Steps:

  1. Choose value of K
  2. Compute distances from test point
  3. Select K nearest neighbors
  4. Take majority vote
  5. Assign class

πŸ”Ή Role of K

The value of:

KK

controls model complexity.


Small K

  • sensitive to noise
  • overfitting

Large K

  • smoother decision boundary
  • underfitting

πŸ”Ή Hyperparameter Tuning

Choosing the best K is called:

Hyperparameter Tuning\textbf{Hyperparameter Tuning}

πŸ”Ή Grid Search

Grid Search systematically tests multiple K values and selects the one with best performance.


πŸ”Ή Cross Validation

Cross Validation divides data into multiple folds:

  • training folds
  • validation folds

Average accuracy is computed across folds.


πŸ”Ή Evaluation Metrics


Accuracy

Accuracy=CorrectPredictionsTotalPredictions Accuracy= \frac{Correct Predictions}{Total Predictions}

πŸ“Š Sample Dataset

A synthetic binary classification dataset is generated using:

make_classification()

Features:

  • 2 input features
  • 2 output classes

πŸ“‹ Algorithm

  1. Import libraries
  2. Generate sample data
  3. Split dataset
  4. Normalize features
  5. Define K values
  6. Apply GridSearchCV
  7. Train KNN model
  8. Find best K
  9. Evaluate accuracy
  10. Plot K vs Accuracy

πŸ’» Program

# ============================================
# KNN CLASSIFICATION USING GRID SEARCH
# ============================================

# --------------------------------------------
# Step 1: Import Libraries
# --------------------------------------------

import numpy as np
import matplotlib.pyplot as plt

from sklearn.datasets import make_classification

from sklearn.model_selection import train_test_split
from sklearn.model_selection import GridSearchCV

from sklearn.preprocessing import StandardScaler

from sklearn.neighbors import KNeighborsClassifier

from sklearn.metrics import accuracy_score
from sklearn.metrics import confusion_matrix
from sklearn.metrics import classification_report

# --------------------------------------------
# Step 2: Generate Sample Dataset
# --------------------------------------------

X, y = make_classification(
    n_samples=200,
    n_features=2,
    n_redundant=0,
    n_informative=2,
    n_classes = 2,
    n_clusters_per_class=1,
    random_state=42
)

print("\nShape of X:", X.shape)
print("Shape of y:", y.shape)
print(X[:10][:10])
print(y[:10])

# --------------------------------------------
# Step 3: Split Dataset
# --------------------------------------------

X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.2,
    random_state=42
)

# --------------------------------------------
# Step 4: Feature Scaling
# --------------------------------------------

scaler = StandardScaler()

X_train = scaler.fit_transform(X_train)

X_test = scaler.transform(X_test)

# --------------------------------------------
# Step 5: Define Model
# --------------------------------------------

knn = KNeighborsClassifier()

# --------------------------------------------
# Step 6: Define Grid Search Parameters
# --------------------------------------------

param_grid = {
    'n_neighbors': list(range(1, 21))
}

# --------------------------------------------
# Step 7: Apply Grid Search
# --------------------------------------------

grid = GridSearchCV(
    estimator=knn,
    param_grid=param_grid,
    cv=5,
    scoring='accuracy'
)

# Train model
grid.fit(X_train, y_train)

# --------------------------------------------
# Step 8: Best Value of K
# --------------------------------------------

best_k = grid.best_params_['n_neighbors']

print("\nBEST VALUE OF K\n")

print("Best K:", best_k)

print("Best Cross Validation Accuracy:",
      round(grid.best_score_, 4))

# --------------------------------------------
# Step 9: Train Final Model
# --------------------------------------------

best_model = grid.best_estimator_

# Predictions
y_pred = best_model.predict(X_test)

# --------------------------------------------
# Step 10: Evaluation
# --------------------------------------------

accuracy = accuracy_score(y_test, y_pred)

print("\nMODEL EVALUATION\n")

print("Accuracy:", round(accuracy, 4))

print("\nConfusion Matrix\n")

print(confusion_matrix(y_test, y_pred))

print("\nClassification Report\n")

print(classification_report(y_test, y_pred))

# --------------------------------------------
# Step 11: Plot K vs Accuracy
# --------------------------------------------

k_values = list(range(1, 21))

mean_scores = grid.cv_results_['mean_test_score']

plt.figure(figsize=(8,5))

plt.plot(k_values,
         mean_scores,
         marker='o')

plt.xlabel("K Value")

plt.ylabel("Cross Validation Accuracy")

plt.title("Finding Optimal K using Grid Search")

plt.grid(True)

plt.show()

# --------------------------------------------
# Step 12: Visualization of Classification
# --------------------------------------------

plt.figure(figsize=(8,6))

plt.scatter(
    X_test[:,0],
    X_test[:,1],
    c=y_pred,
    cmap='bwr',
    edgecolors='k'
)

plt.title("KNN Classification Results")

plt.xlabel("Feature 1")

plt.ylabel("Feature 2")

plt.grid(True)

plt.show()

# ============================================
# END OF PROGRAM
# ============================================

πŸ“Š Sample Output

Shape of X: (200, 2)
Shape of y: (200,)
[[-0.87292898  0.013042  ]
 [ 1.31293463  2.77053357]
 [ 2.34042818  2.42099601]
 [ 2.29454774 -0.40438019]
 [ 0.94410516  0.4772409 ]
 [-0.11959689  0.50891314]
 [ 0.1510847   0.81007677]
 [-0.00745441 -0.45284256]
 [-1.25396925  0.06769236]
 [-0.24392415  1.19979806]]
[1 1 1 1 1 0 1 0 0 0]

BEST VALUE OF K

Best K: 16
Best Cross Validation Accuracy: 0.8688

MODEL EVALUATION

Accuracy: 0.875

Confusion Matrix

[[19  4]
 [ 1 16]]

Classification Report

              precision    recall  f1-score   support

           0       0.95      0.83      0.88        23
           1       0.80      0.94      0.86        17

    accuracy                           0.88        40
   macro avg       0.88      0.88      0.87        40
weighted avg       0.89      0.88      0.88        40



πŸ“ˆ Graph Interpretation


πŸ”Ή K vs Accuracy Graph

  • X-axis → K values
  • Y-axis → Cross-validation accuracy

Highest point indicates:

Optimal K\boxed{ \text{Optimal K} }

πŸ”Ή Classification Plot

  • Red/Blue points represent predicted classes
  • Shows separation learned by KNN

πŸ” Interpretation

ObservationMeaning
Small K    Overfitting
Large K    Underfitting
Optimal K    Best generalization

πŸ“Œ Advantages of Grid Search

✅ Automated hyperparameter tuning
✅ More reliable than manual tuning
✅ Uses cross-validation


πŸ“Œ Limitations

❌ Computationally expensive
❌ Slow for large datasets


✅ Result

The optimal value of K was successfully determined using Grid Search and the KNN classifier achieved good classification accuracy.

Grid Search with Cross Validation effectively identifies the best value of K for KNN classification and improves model performance.

Comments

Popular posts from this blog

Machine Learning Lab PCCSL508 Semester 5 KTU CS 2024 Scheme manual - Dr Binu V P

Lab Assignment-2

Lab Assignment-1