Image Classification using K-Nearest Neighbors (KNN) on Fashion MNIST

 

Experiment

Title

Image Classification using K-Nearest Neighbors (KNN) on Fashion MNIST


๐ŸŽฏ Objective

  • To implement KNN for multi-class image classification
  • To experiment with different values of K
  • To analyze impact on:
    • Accuracy
    • Computational cost

๐Ÿ“Š Dataset: Fashion MNIST

  • 70,000 grayscale images (28×28 pixels)
  • 10 classes:
LabelClass
0T-shirt/top
1Trouser
2Pullover
3Dress
4Coat
5Sandal
6Shirt
7Sneaker
8Bag
9Ankle boot
learn more about fashion MNIT data set

⚙️ Steps

  1. Load dataset
  2. Preprocess (flatten + normalize)
  3. Train KNN
  4. Evaluate performance
  5. Compare different K values

๐Ÿ’ป Complete Python Program

# ------------------------------- # 1. Import Libraries # ------------------------------- import numpy as np from tensorflow.keras.datasets import fashion_mnist from sklearn.neighbors import KNeighborsClassifier from sklearn.metrics import accuracy_score, classification_report from sklearn.preprocessing import StandardScaler # ------------------------------- # 2. Load Dataset # ------------------------------- (X_train, y_train), (X_test, y_test) = fashion_mnist.load_data() # ------------------------------- # 3. Preprocessing # ------------------------------- # Flatten images (28x28 → 784) X_train = X_train.reshape(len(X_train), -1) X_test = X_test.reshape(len(X_test), -1) # Normalize pixel values X_train = X_train / 255.0 X_test = X_test / 255.0 # ------------------------------- # 4. Experiment with different K # ------------------------------- k_values = [1, 3, 5] for k in k_values: print(f"\n===== K = {k} =====") knn = KNeighborsClassifier(n_neighbors=k, n_jobs=-1) knn.fit(X_train[:10000], y_train[:10000]) # subset for speed y_pred = knn.predict(X_test[:2000]) acc = accuracy_score(y_test[:2000], y_pred) print("Accuracy:", acc) print(classification_report(y_test[:2000], y_pred))

๐Ÿ“ˆ Results 

===== K = 1 =====
Accuracy: 0.8075
              precision    recall  f1-score   support

           0       0.75      0.78      0.76       200
           1       0.99      0.96      0.97       203
           2       0.70      0.73      0.72       214
           3       0.83      0.82      0.83       190
           4       0.73      0.69      0.71       219
           5       0.97      0.78      0.86       195
           6       0.55      0.62      0.59       197
           7       0.81      0.92      0.86       200
           8       0.97      0.89      0.93       194
           9       0.86      0.91      0.89       188

    accuracy                           0.81      2000
   macro avg         0.82      0.81      0.81      2000
weighted avg       0.82      0.81      0.81      2000


===== K = 3 =====
Accuracy: 0.8195
              precision    recall  f1-score   support

           0       0.68      0.83      0.75       200
           1       0.99      0.95      0.97       203
           2       0.71      0.79      0.75       214
           3       0.89      0.83      0.86       190
           4       0.78      0.69      0.73       219
           5       0.97      0.79      0.88       195
           6       0.57      0.54      0.56       197
           7       0.86      0.93      0.89       200
           8       0.98      0.90      0.94       194
           9       0.86      0.96      0.91       188

    accuracy                           0.82      2000
   macro avg         0.83      0.82      0.82      2000
weighted avg       0.83      0.82      0.82      2000


===== K = 5 =====
Accuracy: 0.8225
              precision    recall  f1-score   support

           0       0.73      0.83      0.78       200
           1       0.99      0.94      0.97       203
           2       0.72      0.78      0.75       214
           3       0.85      0.85      0.85       190
           4       0.77      0.69      0.73       219
           5       0.98      0.78      0.87       195
           6       0.57      0.56      0.57       197
           7       0.85      0.93      0.89       200
           8       0.97      0.92      0.94       194
           9       0.86      0.96      0.91       188

    accuracy                           0.82      2000
   macro avg         0.83      0.82      0.82      2000
weighted avg       0.83      0.82      0.82      2000

๐Ÿ” Observations

๐Ÿ”น K = 1

  • Very sensitive to noise
  • Overfitting

๐Ÿ”น K = 3

  • Better generalization
  • Balanced performance

๐Ÿ”น K = 5

  • More stable
  • Slight smoothing

๐Ÿ“Š Visualization (Optional)

import matplotlib.pyplot as plt

# Display smaller and clearer images
for i in range(5):
    plt.figure(figsize=(2,2))  # smaller figure size
    plt.imshow(X_test[i].reshape(28,28), cmap='gray', interpolation='nearest')
    plt.title(f"True: {y_test[i]}")
    plt.axis('off')  # remove axes
    plt.show()

๐Ÿงช Lab Tasks

Task 1

Try larger K:

k_values = [7, 9]

Task 2

Use full dataset → observe time


Task 3

Use weighted KNN:

KNeighborsClassifier(weights='distance')

Results

  • KNN works for image classification but:
    • Computationally expensive
  • Optimal K improves:
    • Accuracy
    • Stability

Comments

Popular posts from this blog

Machine Learning Lab PCCSL508 Semester 5 KTU CS 2024 Scheme manual - Dr Binu V P

Lab Assignment-2

Lab Assignment-1