Lab Assignment - 7

 

Assignment-7 week 1 September

Learning Objective: Learn KNN classifier


1.Problem Statement

A college wants to classify students into two categories, Pass and Fail, based on their Study Hours and Attendance Percentage.

The following data is collected from 12 students.

StudentStudy HoursAttendance (%)Result
S12.055Fail
S22.560Fail
S33.058Fail
S43.565Fail
S54.070Pass
S64.572Pass
S75.075Pass
S85.578Pass
S96.080Pass
S106.582Pass
S117.085Pass
S123.075Pass

A new student has the following details:

Study HoursAttendance (%)
4.268

Tasks

  1. Represent the above data using a Pandas DataFrame.
  2. Separate the features (Study Hours, Attendance) and the target variable (Result).
  3. Implement the KNN classification algorithm manually for the new student:
    • Calculate the Euclidean distance from the new student to every training sample.
    • Arrange the samples in ascending order of distance.
    • Select the K nearest neighbors, taking K = 3.
    • Identify the class of each of the 3 neighbors.
    • Use majority voting to predict the result of the new student.
  4. Implement the same classification using KNeighborsClassifier from Scikit-learn.
  5. Compare the manually obtained prediction with the Scikit-learn prediction.
  6. Repeat the experiment for K = 5 and K = 7 and observe whether the prediction changes.
  7. Plot the training samples using a scatter plot and mark the new student's position.

Additional Tasks

  1. Standardize the two features and repeat the KNN classification. Compare the result with the unscaled data.
  2. Explain why feature scaling is important in KNN.
  3. Calculate the classification accuracy using a suitable train-test split.
  4. Construct a confusion matrix for the test predictions.
  5. Try different values of K = 1, 3, 5, 7, 9 and identify the value of K that gives the best accuracy.

2.Problem Statement

Use the Breast Cancer Wisconsin dataset available in Scikit-learn to build a K-Nearest Neighbors (KNN) classifier for predicting whether a breast tumor is malignant or benign.

The objective of the experiment is to understand how Grid Search can be used to systematically find the best value of K (number of nearest neighbors) for a KNN classifier.

Dataset

Use the built-in Scikit-learn dataset:

from sklearn.datasets import load_breast_cancer

data = load_breast_cancer()

The dataset contains diagnostic measurements of breast tumors and a binary target indicating whether the tumor is malignant or benign.

Tasks

  1. Load the Breast Cancer dataset and create a Pandas DataFrame.
  2. Display:
    • Number of samples and features
    • Feature names
    • Target names
    • Class distribution
    • Summary statistics
  3. Separate the features (X) and target (y).
  4. Split the dataset into 80% training data and 20% testing data.
  5. Standardize the features using StandardScaler.
  6. Create a KNN classifier using KNeighborsClassifier.
  7. Use GridSearchCV to find the best value of K from:
    K = 1, 3, 5, 7, 9, 11, 13, 15, 17, 19
  1. Use 5-fold cross-validation during grid search.
  2. Display:
    • Best value of K
    • Best cross-validation accuracy
    • Accuracy of the best model on the test data
  3. Display the confusion matrix, classification report, precision, recall and F1-score for the best KNN model.
  4. Plot K value vs. cross-validation accuracy and identify the best K graphically.
  5. Compare the performance 

3.Problem Statement

Implement the K-Nearest Neighbors (KNN) algorithm for image classification using the Fashion MNIST dataset. Experiment with different values of K and analyze their impact on model performance.
Tasks:
● Load and preprocess the Fashion MNIST dataset.
● Implement KNN for multi-class classification.
● Experiment with different values of K and evaluate performance.
● Discuss the impact of different K values on model accuracy and computational
efficiency.



Comments

Popular posts from this blog

Machine Learning Lab PCCSL508 Semester 5 KTU CS 2024 Scheme manual - Dr Binu V P

Lab Assignment-2

Lab Assignment-1