Linear SVM Classification on Iris Dataset with Decision Boundary Visualization

 

Experiment 

Linear SVM Classification on Iris Dataset with Decision Boundary Visualization


🎯 Objective

To implement a Linear Support Vector Machine (SVM) for binary classification on the Iris dataset and visualize the decision boundary and margin.


πŸ“š Theory

A Linear SVM finds a hyperplane:

w⋅x+b=0w \cdot x + b = 0

that separates two classes with maximum margin.

πŸ”Ή Margin Concept

Margin=2∥w∥\text{Margin} = \frac{2}{\|w\|}

  • Margin = distance between two supporting hyperplanes
  • Larger margin → better generalization
  • Defined by support vectors

πŸ“Š Dataset: Iris Dataset

We use the famous Iris dataset.

  • 3 classes: Setosa, Versicolor, Virginica
  • For simplicity → convert to binary classification:
    • Setosa → 1
    • Non-Setosa → 0

We use only 2 features for visualization:

  • Sepal Length
  • Sepal Width

πŸ”¬ Algorithm / Procedure

  1. Load dataset
  2. Convert into binary classification
  3. Select two features
  4. Train Linear SVM
  5. Plot decision boundary
  6. Plot margin and support vectors

πŸ’» Program

import numpy as np import matplotlib.pyplot as plt from sklearn import svm, datasets # Load dataset iris = datasets.load_iris() X = iris.data[:, :2] # take first 2 features y = iris.target # Convert to binary classification (Setosa vs Non-Setosa) y = np.where(y == 0, 1, -1) # Train Linear SVM model = svm.SVC(kernel='linear', C=1) model.fit(X, y) # Plot data points plt.scatter(X[:, 0], X[:, 1], c=y) # Plot decision boundary ax = plt.gca() xlim = ax.get_xlim() ylim = ax.get_ylim() xx = np.linspace(xlim[0], xlim[1], 100) yy = np.linspace(ylim[0], ylim[1], 100) YY, XX = np.meshgrid(yy, xx) xy = np.vstack([XX.ravel(), YY.ravel()]).T Z = model.decision_function(xy).reshape(XX.shape) # Decision boundary ax.contour(XX, YY, Z, levels=[0]) # Margins ax.contour(XX, YY, Z, levels=[-1, 1], linestyles=['--', '--']) # Support vectors ax.scatter(model.support_vectors_[:, 0], model.support_vectors_[:, 1], s=100, facecolors='none', edgecolors='k') plt.xlabel("Sepal Length") plt.ylabel("Sepal Width") plt.title("Linear SVM on Iris Dataset") plt.show()

πŸ“ˆ Output

  • Scatter plot of Iris data
  • Straight decision boundary
  • Two margin lines (parallel dashed lines)
  • Support vectors highlighted



Discussion: Margin Concept


πŸ”Ή What is Margin?

Margin is the distance between the decision boundary and the closest data points (support vectors).


πŸ”Ή How is Margin Determined?

  • SVM selects hyperplane such that:
    • Distance to nearest points is maximized
  • Only support vectors influence this margin

πŸ”Ή Key Insight

πŸ‘‰ Margin depends on:

  1. Support vectors only
  2. Weight vector w
  3. Regularization parameter C

πŸ”Ή Effect of Margin on Classification

MarginEffect
Large Margin        Better generalization
Small Margin        Risk of overfitting

πŸ”Ή Role of C in Margin

  • Small C → Larger margin, allows some errors
  • Large C → Smaller margin, fewer errors

Result

The Linear SVM model was successfully applied to the Iris dataset for binary classification.

  • The model produced a linear decision boundary separating Setosa from Non-Setosa classes.
  • The maximum margin hyperplane was observed along with parallel margin boundaries.
  • Support vectors were identified as the critical points defining the margin.
  • The classifier demonstrated effective separation with good generalization capability.

Comments

Popular posts from this blog

Machine Learning Lab PCCSL508 Semester 5 KTU CS 2024 Scheme manual - Dr Binu V P

Lab Assignment-2

Lab Assignment-1