Decision Tree using ID3 Algorithm (Iterative Dichotomiser3 Algorithm )- toy example
Experiment
Decision Tree using ID3 Algorithm (Iterative Dichotomiser3 Algorithm )
๐ฏ Objective
- To implement ID3 algorithm
-
To compute:
- Entropy
- Information Gain
- To construct a decision tree for predicting Play Golf
๐ Dataset
Outlook Windy Play
Sunny Weak No
Sunny Strong No
Overcast Weak Yes
Rain Weak Yes
Rain Strong No
๐ Theory
1. Calculate Parent Entropy
The first step is to measure the uncertainty (impurity) of the current dataset before any split occurs.
-
: The proportion of examples in class relative to the total number of examples in
- : The total number of target classes
2. Calculate Weighted Entropy After Split
For a candidate attribute , calculate the Remainder, which is the weighted average entropy of the subsets created by splitting dataset based on .
-
Values(A): The set of all possible values for attribute
-
: The subset of where attribute
- and : Number of elements in the subset and original dataset, respectively
3. Compute Information Gain
The Information Gain for attribute is the difference between the original entropy and the weighted entropy after the split:
4. Attribute Selection (ID3 Algorithm)
The ID3 algorithm computes the Information Gain for every available attribute and selects the one with the highest gain as the decision node.
๐ป Python Program
๐ Expected Output
๐ณ Final Decision Tree (Based on ID3)
๐ Explanation
Root Node:
๐ Outlook (highest information gain)
Branch 1: Overcast
๐ Always → Yes
Branch 1: Sunny
๐ Always → No
Branch 1: Rainy
๐ Split using Windy
๐ Strong→ Yes
๐ Weak → No
๐งช Lab Tasks
Task 1
Manually compute entropy of dataset
Task 2
Verify information gains
๐ Key Insights
| Concept | Observation |
|---|---|
| Entropy | Measures impurity |
| Info Gain | Feature selection |
| ID3 | Greedy algorithm |
๐ง Conclusion
-
ID3 selects:
- Best feature using information gain
-
Builds tree:
- Top-down
-
Works well for:
- Categorical data
Comments
Post a Comment