Hyperparameter Tuning of Neural Network on Fashion MNIST Dataset

 

Experiment

Title:

Hyperparameter Tuning of Neural Network on Fashion MNIST Dataset


๐ŸŽฏ Objective

  • To implement a neural network on Fashion MNIST dataset
  • To experiment with:
    • Learning rate
    • Batch size
    • Number of epochs
  • To analyze the impact of hyperparameters on performance

Theory


๐Ÿ”น 1. Fashion MNIST Dataset

  • Dataset of clothing images (10 classes)
  • Image size: 28 × 28 grayscale
  • Classes: T-shirt, Trouser, Dress, etc.

๐Ÿ”น 2. Hyperparameters

Hyperparameters are user-defined settings that control training.


๐Ÿ”ธ Learning Rate (ฮท)

  • Controls step size in weight updates
ValueEffect
Too small    Slow learning
Too large    Unstable training

๐Ÿ”ธ Batch Size

  • Number of samples per update
ValueEffect
Small        Noisy but better generalization
Large        Faster but may overfit

๐Ÿ”ธ Epochs

  • Number of full passes over dataset
ValueEffect
LowUnderfitting
HighOverfitting

๐Ÿ’ป Program


๐Ÿ“Œ Step 1: Import Libraries

import numpy as np import matplotlib.pyplot as plt import time from tensorflow.keras.datasets import fashion_mnist from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Dense, Flatten, Input from tensorflow.keras.utils import to_categorical from tensorflow.keras.optimizers import Adam

๐Ÿ“Œ Step 2: Load Dataset

(X_train, y_train), (X_test, y_test) = fashion_mnist.load_data()

๐Ÿ“Œ Step 3: Preprocessing

# Normalize X_train = X_train / 255.0 X_test = X_test / 255.0 # One-hot encoding y_train = to_categorical(y_train, 10) y_test = to_categorical(y_test, 10)

๐Ÿ”ง Step 4: Model Function

def create_model(learning_rate): model = Sequential([ Input(shape=(28,28)), Flatten(), Dense(128, activation='relu'), Dense(64, activation='relu'), Dense(10, activation='softmax') ]) optimizer = Adam(learning_rate=learning_rate) model.compile(optimizer=optimizer, loss='categorical_crossentropy', metrics=['accuracy']) return model

๐Ÿงช Step 5: Hyperparameter Experiments

learning_rates = [0.001, 0.01] batch_sizes = [32, 128] epochs_list = [5, 10] results = []

๐Ÿ” Run Experiments

for lr in learning_rates: for batch in batch_sizes: for epochs in epochs_list: print(f"\nLR={lr}, Batch={batch}, Epochs={epochs}") model = create_model(lr) start = time.time() history = model.fit(X_train, y_train, epochs=epochs, batch_size=batch, validation_split=0.2, verbose=0) end = time.time() loss, acc = model.evaluate(X_test, y_test, verbose=0) results.append({ 'lr': lr, 'batch': batch, 'epochs': epochs, 'accuracy': acc, 'time': end - start }) print(f"Accuracy: {acc:.4f}, Time: {end-start:.2f}s")

๐Ÿ“Š Step 6: Display Results

for r in results: print(r)

๐Ÿ“Š Sample Results Table

{'lr': 0.001, 'batch': 32, 'epochs': 5, 'accuracy': 0.8722000122070312, 'time': 39.189568758010864} {'lr': 0.001, 'batch': 32, 'epochs': 10, 'accuracy': 0.8784999847412109, 'time': 65.96800518035889} {'lr': 0.001, 'batch': 128, 'epochs': 5, 'accuracy': 0.8669000267982483, 'time': 13.056512594223022} {'lr': 0.001, 'batch': 128, 'epochs': 10, 'accuracy': 0.8784999847412109, 'time': 23.549450874328613} {'lr': 0.01, 'batch': 32, 'epochs': 5, 'accuracy': 0.8490999937057495, 'time': 38.983253717422485} {'lr': 0.01, 'batch': 32, 'epochs': 10, 'accuracy': 0.859499990940094, 'time': 69.030282497406} {'lr': 0.01, 'batch': 128, 'epochs': 5, 'accuracy': 0.8585000038146973, 'time': 13.016467571258545} {'lr': 0.01, 'batch': 128, 'epochs': 10, 'accuracy': 0.8626000285148621, 'time': 26.107290744781494}

๐Ÿ“ˆ Optional Visualization

accuracies = [r['accuracy'] for r in results] plt.plot(accuracies, marker='o') plt.title("Accuracy Across Experiments") plt.xlabel("Experiment No.") plt.ylabel("Accuracy") plt.show()



๐Ÿ” Observations


✔ Learning Rate

  • 0.001 → stable, better accuracy
  • 0.01 → faster but unstable

✔ Batch Size

  • 32 → better generalization
  • 128 → faster training

✔ Epochs

  • 5 epochs → underfitting
  • 10 epochs → better learning


LR    Batch    Epochs    Accuracy    Observation
0.001    32    10    High    Best
0.01    128    5    Low    Unstable
0.001    128    10    Medium    Balanced

Result

  • Best performance achieved with:
    • Moderate learning rate
    • Smaller batch size
    • Sufficient epochs

Comments

Popular posts from this blog

Machine Learning Lab PCCSL508 Semester 5 KTU CS 2024 Scheme manual - Dr Binu V P

Lab Assignment-2

Lab Assignment-1