Manual Implementation of Gradient Boosting Regression using Sample Data

 

Experiment

Manual Implementation of Gradient Boosting Regression using Sample Data

Aim

To implement Gradient Boosting Regression manually without using ensemble learning libraries and to study the sequential learning process using sample data.


Objective

  1. To understand the concept of ensemble learning.
  2. To understand the concept of boosting.
  3. To understand the working of Gradient Boosting Regression.
  4. To compute residual errors manually.
  5. To implement weak learners without using machine learning ensemble libraries.
  6. To visualize prediction improvement using boosting.

Theory

Ensemble Learning

Ensemble learning is a machine learning technique in which multiple models are combined to improve predictive performance. Instead of relying on a single model, several weak learners are combined to form a stronger predictive model.

Ensemble methods are mainly classified into:

  1. Bagging
  2. Boosting

The main objective of ensemble learning is to improve prediction accuracy and reduce errors by combining multiple models.


Boosting

Boosting is an ensemble learning technique in which multiple weak learners are trained sequentially. Each subsequent learner attempts to reduce the errors made by previous learners.

Working principle:

Weak Learner 1
↓ Error Correction
Weak Learner 2
↓ Error Correction
Weak Learner 3
↓ Error Correction
Final Strong Learner

Unlike bagging where all models are built independently, boosting builds models one after another.


Gradient Boosting Regression

Gradient Boosting builds an additive model by sequentially fitting weak learners to the residual errors.

The final prediction model is:

F(x)=F0(x)+η∑m=1MTm(x)F(x)=F_0(x)+\eta\sum_{m=1}^{M}T_m(x)

Where:

  • F(x)F(x) = Final prediction
  • F0(x)F_0(x) = Initial prediction
  • Tm(x)T_m(x) = Prediction of weak learner mm
  • η\eta = Learning rate
  • MM = Number of weak learners

Initial Prediction

The algorithm starts by predicting the mean of the target values.

F0(x)=1N∑i=1NyiF_0(x)=\frac{1}{N}\sum_{i=1}^{N}y_i

Where:

  • yiy_i = Actual target value
  • NN = Number of samples

Residual Error

Residuals are computed as the difference between actual and predicted values.

ri=yi−F(xi)

Where:

  • rir_i = Residual error
  • yiy_i = Actual output
  • F(xi)F(x_i) = Current prediction

Weak Learner Training

Each weak learner is trained on residual values.

Tm(x)=Fit(X,r)T_m(x)=\text{Fit}(X,r)

where:

  • Tm(x)T_m(x) represents the weak learner prediction
  • rr represents residual errors

Prediction Update

After obtaining the weak learner output, predictions are updated using:

Fm(x)=Fm−1(x)+ηTm(x)F_m(x)=F_{m-1}(x)+\eta T_m(x)

Where:

  • Fm(x)F_m(x) = Updated prediction
  • Fm−1(x)F_{m-1}(x) = Previous prediction
  • Tm(x)T_m(x)= Weak learner output
  • η\eta = Learning rate

Decision Stump

A decision stump is a one-level decision tree with a single split.

Example:

                  X ≤ Threshold
/ \
/ \
Predict A Predict B

Decision stumps act as weak learners in Gradient Boosting.


Algorithm

Step 1: Import required libraries.

Step 2: Create sample dataset.

Step 3: Compute initial prediction using mean value.

Step 4: Initialize prediction values.

Step 5: Compute residual errors.

Step 6: Find the best threshold value for decision stump creation.

Step 7: Generate stump predictions.

Step 8: Update prediction values.

Step 9: Repeat the process for multiple iterations.

Step 10: Display final predictions.

Step 11: Plot prediction curve.


Program

import numpy as np
import matplotlib.pyplot as plt


# Step 1: Create Dataset

X=np.array([1,2,3,4,5,6,7,8])
y=np.array([20,25,35,45,60,80,85,90])
print("Input Values")
print(X)
print("\nTarget Values")
print(y)


# Step 2: Initial Prediction

F0=np.mean(y)
print("\nInitial Prediction")
print(round(F0,2))

pred=np.full(len(y),F0)

learning_rate=0.5

n_estimators=5


# Step 3: Gradient Boosting Process

for t in range(n_estimators):
    print("\nIteration:",t+1)
    residual=y-pred
    print("\nResidual Values")
    print(np.round(residual,2))
    best_error=np.inf
    for threshold in X:
       
        left=np.mean(residual[X<=threshold])
        right=np.mean(residual[X>threshold])
        print("\nThreshold:",threshold)
        print("Left:",round(left,2))
        print("Right:",round(right,2))
        stump_pred=np.where(
            X<=threshold,
            left,
            right
        )

        error=np.sum(
            (
            residual-
            stump_pred
            )**2
        )

        if error<best_error:

            best_error=error
            best_threshold=threshold
            best_left=left
            best_right=right


    stump_pred=np.where(
        X<=best_threshold,
        best_left,
        best_right
    )

    print("\nBest Threshold:",best_threshold)
    print("Best Left:",round(best_left,2))
    print("Best Right:",round(best_right,2))
    print("Best Error:",round(best_error,2))
    print("\nStump Prediction")
    print(np.round(stump_pred,2))
    pred=pred+(learning_rate*stump_pred)
    print("\nUpdated Prediction")
    print(np.round(pred,2))

# Step 4: Final Prediction

print(
    "\nFinal Prediction"
)

print( np.round( pred, 2 ))


# Step 5: Visualization

plt.figure(figsize=(10,6))

plt.scatter(
    X,
    y,
    s=100,
    label='Actual Data'
)

plt.plot(
    X,
    pred,
    linewidth=3,
    label='Boosted Prediction'
)

plt.xlabel(
    "Hours Studied"
)

plt.ylabel(
    "Marks"
)

plt.title(
    "Manual Gradient Boosting Regression"
)

plt.grid()

plt.legend()

plt.show()

Output


Threshold: 1 Left: -23.12 Right: 3.3 Threshold: 2 Left: -20.62 Right: 6.88 Threshold: 3 Left: -16.46 Right: 9.88 Threshold: 4 Left: -11.88 Right: 11.88 Threshold: 5 Left: -10.88 Right: 18.12 Threshold: 6 Left: -6.88 Right: 20.62 Threshold: 7 Left: -3.3 Right: 23.12 Threshold: 8 Left: 0.0 Right: nan Best Threshold: 5 Best Left: -10.88 Best Right: 18.12 Best Error: 438.75 Stump Prediction [-10.88 -10.88 -10.88 -10.88 -10.88 18.12 18.12 18.12] Updated Prediction [37.69 37.69 37.69 37.69 61.44 75.94 75.94 75.94] Iteration: 3 Residual Values [-17.69 -12.69 -2.69 7.31 -1.44 4.06 9.06 14.06] Threshold: 1 Left: -17.69 Right: 2.53 Threshold: 2 Left: -15.19 Right: 5.06 Threshold: 3 Left: -11.02 Right: 6.61 Threshold: 4 Left: -6.44 Right: 6.44 Threshold: 5 Left: -5.44 Right: 9.06 Threshold: 6 Left: -3.85 Right: 11.56 Threshold: 7 Left: -2.01 Right: 14.06 Threshold: 8 Left: 0.0 Right: nan Best Threshold: 2 Best Left: -15.19 Best Right: 5.06 Best Error: 217.88 Stump Prediction [-15.19 -15.19 5.06 5.06 5.06 5.06 5.06 5.06] Updated Prediction [30.09 30.09 40.22 40.22 63.97 78.47 78.47 78.47] Iteration: 4 Residual Values [-10.09 -5.09 -5.22 4.78 -3.97 1.53 6.53 11.53] Threshold: 1 Left: -10.09 Right: 1.44 Threshold: 2 Left: -7.59 Right: 2.53 Threshold: 3 Left: -6.8 Right: 4.08 Threshold: 4 Left: -3.91 Right: 3.91 Threshold: 5 Left: -3.92 Right: 6.53 Threshold: 6 Left: -3.01 Right: 9.03 Threshold: 7 Left: -1.65 Right: 11.53 Threshold: 8 Left: 0.0 Right: nan Best Threshold: 3 Best Left: -6.8 Best Right: 4.08 Best Error: 149.56 Stump Prediction [-6.8 -6.8 -6.8 4.08 4.08 4.08 4.08 4.08] Updated Prediction [26.69 26.69 36.82 42.26 66.01 80.51 80.51 80.51] Iteration: 5 Residual Values [-6.69 -1.69 -1.82 2.74 -6.01 -0.51 4.49 9.49] Threshold: 1 Left: -6.69 Right: 0.96 Threshold: 2 Left: -4.19 Right: 1.4 Threshold: 3 Left: -3.4 Right: 2.04 Threshold: 4 Left: -1.87 Right: 1.87 Threshold: 5 Left: -2.69 Right: 4.49 Threshold: 6 Left: -2.33 Right: 6.99 Threshold: 7 Left: -1.36 Right: 9.49 Threshold: 8 Left: -0.0 Right: nan Best Threshold: 6 Best Left: -2.33 Best Right: 6.99 Best Error: 74.77 Stump Prediction [-2.33 -2.33 -2.33 -2.33 -2.33 -2.33 6.99 6.99] Updated Prediction [25.53 25.53 35.65 41.09 64.84 79.34 84. 84. ] Final Prediction [25.53 25.53 35.65 41.09 64.84 79.34 84. 84. ]


Graph

The graph contains:

  • Scatter points representing actual data
  • Prediction curve showing Gradient Boosting output
  • Gradual improvement of prediction after every iteration

Students can observe how prediction values move closer to actual outputs after each boosting iteration.


Result

 Gradient Boosting Regression was implemented manually without using ensemble learning libraries and the prediction curve was visualized successfully.

Comments

Popular posts from this blog

Machine Learning Lab PCCSL508 Semester 5 KTU CS 2024 Scheme manual - Dr Binu V P

Lab Assignment-2

Lab Assignment-1