Comparison of Activation Functions (Sigmoid, ReLU, Tanh) using Neural Network on MNIST Dataset

 

Experiment

Comparison of Activation Functions (Sigmoid, ReLU, Tanh) using Neural Network on MNIST Dataset


๐ŸŽฏ Objective

  • To implement neural networks using different activation functions
  • To compare:
    • Training performance
    • Convergence speed
    • Classification accuracy

Theory


๐Ÿ”น 1. MNIST Dataset

  • Contains handwritten digit images (0–9)
  • Each image: 28 × 28 pixels (784 features)
  • Total classes: 10

๐Ÿ”น 2. Activation Functions


๐Ÿ”ธ Sigmoid

ฯƒ(x)=11+e−x\sigma(x) = \frac{1}{1 + e^{-x}}
  • Output range: (0,1)
  • Problem: Vanishing gradient

๐Ÿ”ธ Tanh

tanh⁡(x)=ex−e−xex+e−x\tanh(x) = \frac{e^x - e^{-x}}{e^x + e^{-x}}
  • Output range: (-1,1)
  • Better than sigmoid but still vanishing gradient

๐Ÿ”ธ ReLU

f(x)=max⁡(0,x)f(x) = \max(0, x)
  • Fast convergence
  • Most widely used

๐Ÿ”น 3. Comparison Concept

Function        Speed    Gradient Issue
Sigmoid        Slow        Yes
Tanh        Medium        Less
ReLU            Fast        No (mostly)

๐Ÿ’ป Program


๐Ÿ“Œ Step 1: Import Libraries

import numpy as np import matplotlib.pyplot as plt import time from tensorflow.keras.datasets import mnist from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Dense, Input, Flatten from tensorflow.keras.utils import to_categorical

๐Ÿ“Œ Step 2: Load Dataset

(X_train, y_train), (X_test, y_test) = mnist.load_data()

๐Ÿ“Œ Step 3: Preprocessing

# Normalize X_train = X_train / 255.0 X_test = X_test / 255.0 # One-hot encoding y_train = to_categorical(y_train, 10) y_test = to_categorical(y_test, 10)

๐Ÿ”ง Step 4: Model Function

def create_model(activation): model = Sequential([ Input(shape=(28,28)), Flatten(), Dense(128, activation=activation), Dense(64, activation=activation), Dense(10, activation='softmax') ]) model.compile(optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy']) return model

๐Ÿงช Step 5: Train Models with Different Activations

activations = ['sigmoid', 'tanh', 'relu'] results = {} histories = {} for act in activations: print(f"\nTraining with {act} activation") model = create_model(act) start = time.time() history = model.fit(X_train, y_train, epochs=5, batch_size=128, validation_split=0.2, verbose=1) end = time.time() loss, acc = model.evaluate(X_test, y_test, verbose=0) results[act] = { 'accuracy': acc, 'time': end - start } histories[act] = history

๐Ÿ“Š Step 6: Compare Accuracy

for act in results: print(f"{act} -> Accuracy: {results[act]['accuracy']:.4f}, Time: {results[act]['time']:.2f}s")

Output


sigmoid -> Accuracy: 0.9542, Time: 18.48s tanh -> Accuracy: 0.9712, Time: 14.16s relu -> Accuracy: 0.9720, Time: 14.59s

๐Ÿ“ˆ Step 7: Plot Accuracy Curves

for act in activations: plt.plot(histories[act].history['val_accuracy'], label=act) plt.title("Validation Accuracy Comparison") plt.xlabel("Epochs") plt.ylabel("Accuracy") plt.legend() plt.show()




๐Ÿ“ˆ Step 8: Plot Loss Curves

for act in activations: plt.plot(histories[act].history['val_loss'], label=act) plt.title("Validation Loss Comparison") plt.xlabel("Epochs") plt.ylabel("Loss") plt.legend() plt.show()



๐Ÿ” Observations


✔ Sigmoid

  • Slow learning
  • Low accuracy
  • Saturation problem

✔ Tanh

  • Better than sigmoid
  • Faster convergence
  • Still slight gradient issues

✔ ReLU

  • Fastest convergence
  • Highest accuracy
  • Best performance overall

๐Ÿ“Š Sample Result Table

Activation    Accuracy    Training Time    Convergence
Sigmoid    Low    Slow    Poor
Tanh    Medium    Moderate    Better
ReLU    High    Fast    Best

Result

  • ReLU performs best in terms of:
    • Accuracy
    • Speed
  • Sigmoid performs worst due to vanishing gradient

Comments

Popular posts from this blog

Machine Learning Lab PCCSL508 Semester 5 KTU CS 2024 Scheme manual - Dr Binu V P

Lab Assignment-2

Lab Assignment-1