Finding the Optimal Value of K in KNN using Grid Search
Experiment:
π Aim
To implement K-Nearest Neighbors (KNN) Classification and determine the best value of K using Grid Search with Cross Validation.
π― Objectives
- Understand KNN classification
- Generate sample classification data
- Implement KNN classifier
- Use GridSearchCV to find optimal K
- Evaluate classification performance
- Visualize accuracy for different K values
π Theory
πΉ K-Nearest Neighbors (KNN)
KNN is a supervised machine learning algorithm used for:
- classification
- regression
In KNN classification:
- the class of a new point is determined by the majority class among its K nearest neighbors.
πΉ Working of KNN
Steps:
- Choose value of K
- Compute distances from test point
- Select K nearest neighbors
- Take majority vote
- Assign class
πΉ Role of K
The value of:
controls model complexity.
Small K
- sensitive to noise
- overfitting
Large K
- smoother decision boundary
- underfitting
πΉ Hyperparameter Tuning
Choosing the best K is called:
πΉ Grid Search
Grid Search systematically tests multiple K values and selects the one with best performance.
πΉ Cross Validation
Cross Validation divides data into multiple folds:
- training folds
- validation folds
Average accuracy is computed across folds.
πΉ Evaluation Metrics
Accuracy
π Sample Dataset
A synthetic binary classification dataset is generated using:
make_classification()
Features:
- 2 input features
- 2 output classes
π Algorithm
- Import libraries
- Generate sample data
- Split dataset
- Normalize features
- Define K values
- Apply GridSearchCV
- Train KNN model
- Find best K
- Evaluate accuracy
- Plot K vs Accuracy
π» Program
# ============================================# KNN CLASSIFICATION USING GRID SEARCH# ============================================# --------------------------------------------# Step 1: Import Libraries# --------------------------------------------import numpy as npimport matplotlib.pyplot as pltfrom sklearn.datasets import make_classificationfrom sklearn.model_selection import train_test_splitfrom sklearn.model_selection import GridSearchCVfrom sklearn.preprocessing import StandardScalerfrom sklearn.neighbors import KNeighborsClassifierfrom sklearn.metrics import accuracy_scorefrom sklearn.metrics import confusion_matrixfrom sklearn.metrics import classification_report# --------------------------------------------# Step 2: Generate Sample Dataset# --------------------------------------------X, y = make_classification(n_samples=200,n_features=2,n_redundant=0,n_informative=2,n_classes = 2,n_clusters_per_class=1,random_state=42)print("\nShape of X:", X.shape)print("Shape of y:", y.shape)print(X[:10][:10])print(y[:10])# --------------------------------------------# Step 3: Split Dataset# --------------------------------------------X_train, X_test, y_train, y_test = train_test_split(X,y,test_size=0.2,random_state=42)# --------------------------------------------# Step 4: Feature Scaling# --------------------------------------------scaler = StandardScaler()X_train = scaler.fit_transform(X_train)X_test = scaler.transform(X_test)# --------------------------------------------# Step 5: Define Model# --------------------------------------------knn = KNeighborsClassifier()# --------------------------------------------# Step 6: Define Grid Search Parameters# --------------------------------------------param_grid = {'n_neighbors': list(range(1, 21))}# --------------------------------------------# Step 7: Apply Grid Search# --------------------------------------------grid = GridSearchCV(estimator=knn,param_grid=param_grid,cv=5,scoring='accuracy')# Train modelgrid.fit(X_train, y_train)# --------------------------------------------# Step 8: Best Value of K# --------------------------------------------best_k = grid.best_params_['n_neighbors']print("\nBEST VALUE OF K\n")print("Best K:", best_k)print("Best Cross Validation Accuracy:",round(grid.best_score_, 4))# --------------------------------------------# Step 9: Train Final Model# --------------------------------------------best_model = grid.best_estimator_# Predictionsy_pred = best_model.predict(X_test)# --------------------------------------------# Step 10: Evaluation# --------------------------------------------accuracy = accuracy_score(y_test, y_pred)print("\nMODEL EVALUATION\n")print("Accuracy:", round(accuracy, 4))print("\nConfusion Matrix\n")print(confusion_matrix(y_test, y_pred))print("\nClassification Report\n")print(classification_report(y_test, y_pred))# --------------------------------------------# Step 11: Plot K vs Accuracy# --------------------------------------------k_values = list(range(1, 21))mean_scores = grid.cv_results_['mean_test_score']plt.figure(figsize=(8,5))plt.plot(k_values,mean_scores,marker='o')plt.xlabel("K Value")plt.ylabel("Cross Validation Accuracy")plt.title("Finding Optimal K using Grid Search")plt.grid(True)plt.show()# --------------------------------------------# Step 12: Visualization of Classification# --------------------------------------------plt.figure(figsize=(8,6))plt.scatter(X_test[:,0],X_test[:,1],c=y_pred,cmap='bwr',edgecolors='k')plt.title("KNN Classification Results")plt.xlabel("Feature 1")plt.ylabel("Feature 2")plt.grid(True)plt.show()# ============================================# END OF PROGRAM# ============================================
π Sample Output
Shape of X: (200, 2) Shape of y: (200,) [[-0.87292898 0.013042 ] [ 1.31293463 2.77053357] [ 2.34042818 2.42099601] [ 2.29454774 -0.40438019] [ 0.94410516 0.4772409 ] [-0.11959689 0.50891314] [ 0.1510847 0.81007677] [-0.00745441 -0.45284256] [-1.25396925 0.06769236] [-0.24392415 1.19979806]] [1 1 1 1 1 0 1 0 0 0] BEST VALUE OF K Best K: 16 Best Cross Validation Accuracy: 0.8688 MODEL EVALUATION Accuracy: 0.875 Confusion Matrix [[19 4] [ 1 16]] Classification Report precision recall f1-score support 0 0.95 0.83 0.88 23 1 0.80 0.94 0.86 17 accuracy 0.88 40 macro avg 0.88 0.88 0.87 40 weighted avg 0.89 0.88 0.88 40
π Graph Interpretation
πΉ K vs Accuracy Graph
- X-axis → K values
- Y-axis → Cross-validation accuracy
Highest point indicates:
πΉ Classification Plot
- Red/Blue points represent predicted classes
- Shows separation learned by KNN
π Interpretation
| Observation | Meaning |
|---|---|
| Small K | Overfitting |
| Large K | Underfitting |
| Optimal K | Best generalization |
π Advantages of Grid Search
✅ Automated hyperparameter tuning
✅ More reliable than manual tuning
✅ Uses cross-validation
π Limitations
❌ Computationally expensive
❌ Slow for large datasets
✅ Result
The optimal value of K was successfully determined using Grid Search and the KNN classifier achieved good classification accuracy.
Grid Search with Cross Validation effectively identifies the best value of K for KNN classification and improves model performance.


Comments
Post a Comment