Comparison of Logistic Regression and Decision Tree on Adult Income Dataset

 

Experiment

Title

Comparison of Logistic Regression and Decision Tree on Adult Income Dataset


🎯 Objective

  • To predict whether income is:
    • <=50K or >50K
  • To compare:
    • Logistic Regression
    • Decision Tree
  • To evaluate using:
    • Accuracy
    • Precision
    • Recall
    • F1-score

πŸ“Š Dataset: Adult Income

Features include:

  • Age
  • Education
  • Occupation
  • Hours-per-week
  • Capital gain/loss

Target:

  • Income (<=50K, >50K)

⚙️ Steps

  1. Load dataset
  2. Handle missing values
  3. Encode categorical variables
  4. Train models
  5. Evaluate performance
  6. Compare interpretability

πŸ’» Python Program

# ------------------------------- # 1. Import Libraries # ------------------------------- import pandas as pd from sklearn.model_selection import train_test_split from sklearn.preprocessing import LabelEncoder, StandardScaler from sklearn.linear_model import LogisticRegression from sklearn.tree import DecisionTreeClassifier from sklearn.metrics import classification_report, accuracy_score # ------------------------------- # 2. Load Dataset # ------------------------------- # Download from: # https://archive.ics.uci.edu/ml/datasets/adult columns = ['age','workclass','fnlwgt','education','education-num', 'marital-status','occupation','relationship','race', 'sex','capital-gain','capital-loss','hours-per-week', 'native-country','income'] df = pd.read_csv("adult.data", names=columns, na_values=" ?", skipinitialspace=True) # ------------------------------- # 3. Preprocessing # ------------------------------- # Remove missing values df.dropna(inplace=True) # Encode categorical variables le = LabelEncoder() for col in df.select_dtypes(include='object').columns: df[col] = le.fit_transform(df[col]) # Features and target X = df.drop('income', axis=1) y = df['income'] # Scale for Logistic Regression scaler = StandardScaler() X_scaled = scaler.fit_transform(X) # Train-test split X_train, X_test, y_train, y_test = train_test_split( X_scaled, y, test_size=0.2, random_state=42 ) # ------------------------------- # 4. Logistic Regression # ------------------------------- lr = LogisticRegression(max_iter=1000) lr.fit(X_train, y_train) y_pred_lr = lr.predict(X_test) # ------------------------------- # 5. Decision Tree (ID3) # ------------------------------- dt = DecisionTreeClassifier(criterion='entropy', max_depth=5) dt.fit(X_train, y_train) y_pred_dt = dt.predict(X_test) # ------------------------------- # 6. Evaluation # ------------------------------- print("=== Logistic Regression ===") print("Accuracy:", accuracy_score(y_test, y_pred_lr)) print(classification_report(y_test, y_pred_lr)) print("\n=== Decision Tree ===") print("Accuracy:", accuracy_score(y_test, y_pred_dt)) print(classification_report(y_test, y_pred_dt))

πŸ“ˆResults 

=== Logistic Regression === Accuracy: 0.8246583755565792 precision recall f1-score support 0 0.85 0.94 0.89 4942 1 0.71 0.46 0.56 1571 accuracy 0.82 6513 macro avg 0.78 0.70 0.72 6513 weighted avg 0.81 0.82 0.81 6513 === Decision Tree === Accuracy: 0.8490710885920467 precision recall f1-score support 0 0.86 0.95 0.91 4942 1 0.78 0.52 0.62 1571 accuracy 0.85 6513 macro avg 0.82 0.74 0.76 6513 weighted avg 0.84 0.85 0.84 6513

πŸ“Š Observations

πŸ”Ή Logistic Regression

  • Linear model
  • Stable performance
  • Works well with scaled data

πŸ”Ή Decision Tree

  • Non-linear model
  • Captures complex patterns
  • May overfit (especially deep trees)

πŸ“Š Interpretability Comparison

Aspect    Logistic Regression    Decision Tree
Model Type    Linear    Rule-based
Interpretability    Medium    High ✅
Explanation    Coefficients    If-else rules
Visualization    Difficult    Easy

🌳 Decision Tree Interpretation

Example rule:

IF hours-per-week > 40 AND education-num > 10 → Income >50K

πŸ‘‰ Easy to understand


πŸ“ˆ Logistic Regression Interpretation

  • Uses coefficients:
y=w1x1+w2x2+...y = w_1x_1 + w_2x_2 + ...

πŸ‘‰ Harder to interpret for beginners


Key Insights

  • Logistic Regression:
    • Better generalization
  • Decision Tree:
    • Better interpretability

⚖️ When to Use What?

SituationBest Model
Need explainability    Decision Tree
Need accuracy & stability    Logistic Regression
Non-linear data    Decision Tree

πŸ§ͺ Lab Tasks

Task 1

Change tree depth:

max_depth = 3, 10

Task 2

Remove scaling → observe LR performance


Task 3

Plot tree:

from sklearn.tree import plot_tree

Result

  • Logistic Regression:
    • Stable, good baseline
  • Decision Tree:
    • Easy to interpret
  • Trade-off:
    • Accuracy vs interpretability

Comments

Popular posts from this blog

Machine Learning Lab PCCSL508 Semester 5 KTU CS 2024 Scheme manual - Dr Binu V P

Lab Assignment-2

Lab Assignment-1