Lab Assignment - 7
Assignment-7 week 1 September
Learning Objective: Learn KNN classifier
1.Problem Statement
A college wants to classify students into two categories, Pass and Fail, based on their Study Hours and Attendance Percentage.
The following data is collected from 12 students.
| Student | Study Hours | Attendance (%) | Result |
|---|---|---|---|
| S1 | 2.0 | 55 | Fail |
| S2 | 2.5 | 60 | Fail |
| S3 | 3.0 | 58 | Fail |
| S4 | 3.5 | 65 | Fail |
| S5 | 4.0 | 70 | Pass |
| S6 | 4.5 | 72 | Pass |
| S7 | 5.0 | 75 | Pass |
| S8 | 5.5 | 78 | Pass |
| S9 | 6.0 | 80 | Pass |
| S10 | 6.5 | 82 | Pass |
| S11 | 7.0 | 85 | Pass |
| S12 | 3.0 | 75 | Pass |
A new student has the following details:
| Study Hours | Attendance (%) |
|---|---|
| 4.2 | 68 |
Tasks
- Represent the above data using a Pandas DataFrame.
-
Separate the features (
Study Hours,Attendance) and the target variable (Result). -
Implement the KNN classification algorithm manually for the new student:
- Calculate the Euclidean distance from the new student to every training sample.
- Arrange the samples in ascending order of distance.
- Select the K nearest neighbors, taking K = 3.
- Identify the class of each of the 3 neighbors.
- Use majority voting to predict the result of the new student.
-
Implement the same classification using
KNeighborsClassifierfrom Scikit-learn. - Compare the manually obtained prediction with the Scikit-learn prediction.
- Repeat the experiment for K = 5 and K = 7 and observe whether the prediction changes.
- Plot the training samples using a scatter plot and mark the new student's position.
Additional Tasks
- Standardize the two features and repeat the KNN classification. Compare the result with the unscaled data.
- Explain why feature scaling is important in KNN.
- Calculate the classification accuracy using a suitable train-test split.
- Construct a confusion matrix for the test predictions.
- Try different values of K = 1, 3, 5, 7, 9 and identify the value of K that gives the best accuracy.
2.Problem Statement
Use the Breast Cancer Wisconsin dataset available in Scikit-learn to build a K-Nearest Neighbors (KNN) classifier for predicting whether a breast tumor is malignant or benign.
The objective of the experiment is to understand how Grid Search can be used to systematically find the best value of K (number of nearest neighbors) for a KNN classifier.
Dataset
Use the built-in Scikit-learn dataset:
from sklearn.datasets import load_breast_cancer data = load_breast_cancer()
The dataset contains diagnostic measurements of breast tumors and a binary target indicating whether the tumor is malignant or benign.
Tasks
- Load the Breast Cancer dataset and create a Pandas DataFrame.
-
Display:
- Number of samples and features
- Feature names
- Target names
- Class distribution
- Summary statistics
-
Separate the features (
X) and target (y). - Split the dataset into 80% training data and 20% testing data.
-
Standardize the features using
StandardScaler. -
Create a KNN classifier using
KNeighborsClassifier. -
Use GridSearchCV to find the best value of
Kfrom:
K = 1, 3, 5, 7, 9, 11, 13, 15, 17, 19
- Use 5-fold cross-validation during grid search.
-
Display:
- Best value of K
- Best cross-validation accuracy
- Accuracy of the best model on the test data
- Display the confusion matrix, classification report, precision, recall and F1-score for the best KNN model.
- Plot K value vs. cross-validation accuracy and identify the best K graphically.
- Compare the performance
3.Problem Statement
Implement the K-Nearest Neighbors (KNN) algorithm for image classification using the Fashion MNIST dataset. Experiment with different values of K and analyze their impact on model performance.
Tasks:
● Load and preprocess the Fashion MNIST dataset.
● Implement KNN for multi-class classification.
● Experiment with different values of K and evaluate performance.
● Discuss the impact of different K values on model accuracy and computational
efficiency.
Comments
Post a Comment