Customer Segmentation using Decision Tree (ID3) on Online Retail Dataset

 

Experiment

Title

Customer Segmentation using Decision Tree (ID3) on Online Retail Dataset


๐ŸŽฏ Objective

  • To segment customers based on purchasing behavior
  • To implement Decision Tree using ID3 (entropy)
  • To analyze:
    • Tree structure
    • Feature importance

๐Ÿ“Š Dataset: Online Retail

Typical features:

  • Quantity → Number of items purchased
  • UnitPrice → Price per item
  • Country → Customer location
  • InvoiceDate → Purchase date
  • CustomerID

๐ŸŽฏ Target (Segmentation Idea)

We create a simple segmentation:

๐Ÿ‘‰ High Value Customer (1)
๐Ÿ‘‰ Low Value Customer (0)

Based on:

Total Spend=Quantity×UnitPrice\text{Total Spend} = Quantity \times UnitPrice

⚙️ Steps

  1. Load dataset
  2. Clean data
  3. Create features
  4. Define target
  5. Train ID3 model
  6. Visualize tree
  7. Analyze feature importance

๐Ÿ’ป Python Program

# ------------------------------- # 1. Import Libraries # ------------------------------- import pandas as pd import numpy as np from sklearn.tree import DecisionTreeClassifier, plot_tree from sklearn.preprocessing import LabelEncoder import matplotlib.pyplot as plt # ------------------------------- # 2. Load Dataset # ------------------------------- # Download from: # https://archive.ics.uci.edu/ml/datasets/Online+Retail df = pd.read_excel("Online Retail.xlsx") # ------------------------------- # 3. Data Preprocessing # ------------------------------- # Remove missing values df = df.dropna(subset=['CustomerID']) # Remove negative/invalid values df = df[(df['Quantity'] > 0) & (df['UnitPrice'] > 0)] # Create Total Spend df['TotalSpend'] = df['Quantity'] * df['UnitPrice'] # ------------------------------- # 4. Customer-Level Aggregation # ------------------------------- customer_df = df.groupby('CustomerID').agg({ 'Quantity': 'sum', 'UnitPrice': 'mean', 'TotalSpend': 'sum', 'Country': 'first' }).reset_index() # ------------------------------- # 5. Create Target Variable # ------------------------------- # High Value if above median spend threshold = customer_df['TotalSpend'].median() customer_df['HighValue'] = (customer_df['TotalSpend'] > threshold).astype(int) # ------------------------------- # 6. Encode Categorical Data # ------------------------------- le = LabelEncoder() customer_df['Country'] = le.fit_transform(customer_df['Country']) # Features and target X = customer_df[['Quantity', 'UnitPrice', 'Country']] y = customer_df['HighValue'] # ------------------------------- # 7. Train Decision Tree (ID3) # ------------------------------- model = DecisionTreeClassifier(criterion='entropy', max_depth=4) model.fit(X, y) # ------------------------------- # 8. Visualization # ------------------------------- plt.figure(figsize=(20,8)) plot_tree(model, feature_names=X.columns, class_names=['Low','High'], filled=True) plt.title("Customer Segmentation using Decision Tree (ID3)") plt.show() # ------------------------------- # 9. Feature Importance # ------------------------------- importance = pd.Series(model.feature_importances_, index=X.columns) print("\nFeature Importance:\n", importance.sort_values(ascending=False))

๐Ÿ“ˆ  Results

๐Ÿ”น Tree Structure

  • Root node often = Total spending related feature (Quantity / UnitPrice)
  • Splits based on purchasing patterns

๐Ÿ”น Feature Importance 

Feature Importance: Quantity 0.86527 UnitPrice 0.13473 Country 0.00000 dtype: float64







๐ŸŒณ Interpretation of Tree

Example Logic:

If Quantity > threshold: → High Value Customer Else: Check UnitPrice

๐Ÿ” Analysis

๐Ÿ”น How Tree Helps Understand Customer Behavior

1. Spending Patterns

  • High quantity → bulk buyers
  • High price → premium buyers

2. Customer Segmentation

  • Low spenders
  • High spenders

3. Business Insights

PatternInsight
High Quantity    Wholesale customers
High UnitPrice    Premium buyers
Country influence    Regional behavior

๐Ÿ“Š Key Insights

  • Decision trees are:
    • Interpretable
    • Easy to visualize
  • ID3 uses:
    • Entropy
    • Information gain

⚖️ Advantages

  • Easy to explain
  • Works well with categorical + numeric data

❌ Limitations

  • Overfitting
  • Sensitive to noise

๐Ÿงช Lab Tasks

Task 1

Change depth:

max_depth = 3, 5

Task 2

Add more features:

Invoice frequency

Task 3

Compare with:

criterion='gini'

Result

  • Decision Tree (ID3) effectively segments customers
  • Feature importance reveals:
    • Key buying factors
  • Tree structure provides:
    • Clear business rules

Comments

Popular posts from this blog

Machine Learning Lab PCCSL508 Semester 5 KTU CS 2024 Scheme manual - Dr Binu V P

Lab Assignment-2

Lab Assignment-1