Implementation of K-Nearest Neighbors (KNN) Classification Algorithm - Sample data Set

 

Experiment

Title

Implementation of K-Nearest Neighbors (KNN) Classification Algorithm


๐ŸŽฏ Objective

  • To understand the working of KNN classification
  • To classify data points based on nearest neighbors
  • To study the effect of different values of K

๐Ÿ“š Background Theory

๐Ÿ”น KNN Algorithm

  • A non-parametric, instance-based learning algorithm
  • Classifies a new point based on majority vote of nearest neighbors

๐Ÿ”น Distance Measure (Euclidean Distance)

d=(x1−x2)2+(y1−y2)2d = \sqrt{(x_1 - x_2)^2 + (y_1 - y_2)^2}

๐Ÿงฉ Sample Dataset

We classify whether a fruit is:

  • ๐ŸŽ Apple (0)
  • ๐ŸŠ Orange (1)
WeightSizeClass
1507Apple
1707.5Apple
1406.8Apple
1306.5Orange
1206.2Orange
1106Orange

⚙️ Algorithm Steps

  1. Choose value of K
  2. Compute distance from test point to all training points
  3. Select K nearest neighbors
  4. Perform majority voting
  5. Assign class

๐Ÿ’ป Python Program

import numpy as np from collections import Counter # ------------------------------- # 1. Dataset # ------------------------------- # Features: [Weight, Size] X = np.array([ [150, 7], [170, 7.5], [140, 6.8], [130, 6.5], [120, 6.2], [110, 6] ]) # Labels: 0 = Apple, 1 = Orange y = np.array([0, 0, 0, 1, 1, 1]) # ------------------------------- # 2. Distance Function # ------------------------------- def euclidean_distance(a, b): return float(np.sqrt(np.sum((a - b)**2))) # ------------------------------- # 3. KNN Function # ------------------------------- def knn_predict(X, y, test_point, k=3): distances = [] # Compute distances for i in range(len(X)): dist = euclidean_distance(X[i], test_point) distances.append((dist, float(y[i]))) # Sort by distance distances.sort(key=lambda x: x[0]) print("sorted distances") print(distances) # Select K nearest neighbors = distances[:k] # Voting labels = [label for _, label in neighbors] prediction = Counter(labels).most_common(1)[0][0] return prediction, neighbors # ------------------------------- # 4. Test Point # ------------------------------- test_point = np.array([135, 6.7]) print("Test Point:", test_point) prediction, neighbors = knn_predict(X, y, test_point, k=3) # ------------------------------- # 5. Output # ------------------------------- print("Nearest Neighbors:", neighbors) if prediction == 0: print("Predicted Class: Apple") else: print("Predicted Class: Orange")

๐Ÿ“ˆ  Output

Test Point: [135. 6.7] sorted distances [(5.000999900019995, 0.0), (5.0039984012787215, 1.0), (15.002999700059986, 0.0), (15.008331019803634, 1.0), (25.009798079952585, 1.0), (35.009141663285604, 0.0)] Nearest Neighbors: [(5.000999900019995, 0.0), (5.0039984012787215, 1.0), (15.002999700059986, 0.0)] Predicted Class: Apple

๐Ÿ” Observations

  • Classification depends on nearest neighbors
  • Changing K changes prediction
  • Smaller K → sensitive to noise
  • Larger K → smoother decision

๐Ÿงช Lab Tasks

Task 1

Change value of K:

k = 1, 3, 5

Task 2

Change test point:

[115, 6.1]

๐Ÿ“Š Key Insights

K Value    Behavior
Small K    Overfitting
Large K    Underfitting

Result

  • KNN is:
    • Simple
    • Easy to implement
  • No training phase (lazy learner)
  • Works well for:
    • Small datasets

Comments

Popular posts from this blog

Machine Learning Lab PCCSL508 Semester 5 KTU CS 2024 Scheme manual - Dr Binu V P

Lab Assignment-2

Lab Assignment-1