Implementation of XOR Problem using a Multilayer Neural Network with Backpropagation (From Scratch)

 

Experiment

Title

Implementation of XOR Problem using a Multilayer Neural Network with Backpropagation (From Scratch)


🎯 Objective

  • To implement a neural network for solving the XOR problem
  • To understand forward propagation
  • To derive and implement backpropagation
  • To update weights and biases using gradient descent
  • To visualize the learning process and decision boundary

🧠 Theory

1. XOR Problem

The XOR (Exclusive OR) operation gives output 1 only when the two inputs are different.

x₁x₂y
000
011
101
110

The XOR problem is not linearly separable, meaning a single straight line cannot separate the classes.

Hence a single-layer neural network fails.

A hidden layer is required.


2. Network Architecture

Neural network used:

  • Input layer → 2 neurons
  • Hidden layer → 2 neurons
  • Output layer → 1 neuron

Architecture:

                    Input Layer          Hidden Layer         Output Layer




Biases:

  • Hidden neuron 1 → b₁
  • Hidden neuron 2 → b₂
  • Output neuron → b₃

🔹 Forward Propagation


Hidden neuron 1

z1=x1w1+x2w2+b1z_1=x_1w_1+x_2w_2+b_1 h1=σ(z1)h_1=\sigma(z_1)

Hidden neuron 2

z2=x1w3+x2w4+b2z_2=x_1w_3+x_2w_4+b_2 h2=σ(z2)h_2=\sigma(z_2)

Output neuron

z3=h1w5+h2w6+b3z_3=h_1w_5+h_2w_6+b_3 y^=σ(z3)\hat y=\sigma(z_3)

Sigmoid Activation Function

σ(x)=11+e−x\sigma(x) = \frac{1}{1+e^{-x}}

Sigmoid derivative:

σ′(x)=σ(x)(1−σ(x))\sigma'(x) = \sigma(x)(1-\sigma(x))

Error Function

Mean Squared Error:

E=12(y−y^)2E = \frac12(y-\hat y)^2

🔹 Backpropagation

Backpropagation computes gradients using chain rule and propagates error backward.


Output Error

erroro=(y^−y)error_o = (\hat y-y)

Output Delta

δo=(y^−y)y^(1−y^)\delta_o = (\hat y-y) \hat y(1-\hat y)

Hidden Layer Errors

For hidden neuron 1:

errorh1=δow5error_{h1} = \delta_ow_5

For hidden neuron 2:

errorh2=δow6error_{h2} = \delta_ow_6

Hidden Layer Deltas

δh1=errorh1×h1(1−h1)\delta_{h1} = error_{h1} \times h_1(1-h_1)
δh2=errorh2×h2(1−h2)\delta_{h2} = error_{h2} \times h_2(1-h_2)

🔹 Gradient Descent Weight Update Rule

Standard gradient descent rule:

wnew=wold−η∂E∂ww_{new} = w_{old} - \eta \frac{\partial E}{\partial w}

where:

  • η\eta = learning rate

Weight updates:

Output layer:

w5=w5−η∑(h1δo)w_5 = w_5 - \eta \sum(h_1\delta_o)
w6=w6−η∑(h2δo)w_6 = w_6 - \eta \sum(h_2\delta_o)

Bias:

b3=b3−η∑(δo)b_3 = b_3 - \eta\sum(\delta_o)

Hidden layer:

w1=w1−η∑(x1δh1)w_1 = w_1 - \eta \sum(x_1\delta_{h1})
w2=w2−η∑(x2δh1)w_2 = w_2 - \eta \sum(x_2\delta_{h1})
w3=w3−η∑(x1δh2)w_3 = w_3 - \eta \sum(x_1\delta_{h2})
w4=w4−η∑(x2δh2)w_4 = w_4 - \eta \sum(x_2\delta_{h2})

Biases:

b1=b1−η∑(δh1)b_1 = b_1-\eta\sum(\delta_{h1})
b2=b2−η∑(δh2)b_2 = b_2-\eta\sum(\delta_{h2})

💻 Program

import numpy as np
import matplotlib.pyplot as plt

# Sigmoid
def sigmoid(x):
return 1/(1+np.exp(-x))

def sigmoid_derivative(x):
return x*(1-x)

# Input data
x1=np.array([0,0,1,1])
x2=np.array([0,1,0,1])

# XOR output
y=np.array([0,1,1,0])

np.random.seed(0)

# Hidden layer weights
w1=np.random.rand()
w2=np.random.rand()

w3=np.random.rand()
w4=np.random.rand()

b1=np.random.rand()
b2=np.random.rand()

# Output layer weights
w5=np.random.rand()
w6=np.random.rand()

b3=np.random.rand()

lr=0.1
epochs=10000

losses=[]

for epoch in range(epochs):

# ====================
# Forward Propagation
# ====================

z1=x1*w1+x2*w2+b1
h1=sigmoid(z1)

z2=x1*w3+x2*w4+b2
h2=sigmoid(z2)

z3=h1*w5+h2*w6+b3
output=sigmoid(z3)

# Loss
loss=np.mean((y-output)**2)
losses.append(loss)

# ====================
# Backpropagation
# ====================

output_error=(output-y)

output_delta=(
output_error*
sigmoid_derivative(output)
)

hidden1_error=output_delta*w5
hidden2_error=output_delta*w6

hidden1_delta=(
hidden1_error*
sigmoid_derivative(h1)
)

hidden2_delta=(
hidden2_error*
sigmoid_derivative(h2)
)

# ====================
# Weight updates
# ====================

w5 -= lr*np.sum(h1*output_delta)
w6 -= lr*np.sum(h2*output_delta)

b3 -= lr*np.sum(output_delta)

w1 -= lr*np.sum(x1*hidden1_delta)
w2 -= lr*np.sum(x2*hidden1_delta)

w3 -= lr*np.sum(x1*hidden2_delta)
w4 -= lr*np.sum(x2*hidden2_delta)

b1 -= lr*np.sum(hidden1_delta)
b2 -= lr*np.sum(hidden2_delta)


print("Predictions:")
print(np.round(output,3))


# Loss curve
plt.plot(losses)
plt.title("Loss vs Epochs")
plt.xlabel("Epoch")
plt.ylabel("Loss")
plt.show()




📊 Decision Boundary Visualization

def predict(a,b):

h1=sigmoid(a*w1+b*w2+b1)

h2=sigmoid(a*w3+b*w4+b2)

out=sigmoid(h1*w5+h2*w6+b3)

return out


xx,yy=np.meshgrid(
np.linspace(-0.5,1.5,200),
np.linspace(-0.5,1.5,200)
)

Z=np.zeros(xx.shape)

for i in range(xx.shape[0]):
for j in range(xx.shape[1]):
Z[i,j]=predict(xx[i,j],yy[i,j])

plt.contourf(xx,yy,Z,
levels=50,
cmap=plt.cm.coolwarm)

plt.scatter(
x1,
x2,
c=y,
s=100,
edgecolors='black'
)

plt.xlabel("x1")
plt.ylabel("x2")
plt.title("XOR Decision Boundary")

plt.show()




🔍 Observations

  • Loss decreases gradually
  • Network learns XOR correctly
  • Decision boundary becomes non-linear
  • Hidden layer enables separation of XOR classes

🧪 Result

The neural network successfully learned the XOR problem using forward propagation and backpropagation.

Comments

Popular posts from this blog

Machine Learning Lab PCCSL508 Semester 5 KTU CS 2024 Scheme manual - Dr Binu V P

Lab Assignment-2

Lab Assignment-1