Logistic Regression on Pima Indians Diabetes Dataset
Experiment
Logistic Regression on Pima Indians Diabetes Dataset🎯 Aim
To implement Logistic Regression for disease prediction and compare model performance with and without feature scaling.
📘 Theory
Logistic Regression predicts probability using:
- Output ≥ 0.5 → Diabetic(diabetes = 1)
- Output < 0.5 → No disease (0)
📊 Dataset Description
The Pima Indians Diabetes Dataset contains medical attributes:
| Feature | Description |
|---|---|
| Pregnancies | Number of pregnancies |
| Glucose | Blood glucose level |
| BloodPressure | Blood pressure |
| SkinThickness | Skin fold thickness |
| Insulin | Insulin level |
| BMI | Body mass index |
| DiabetesPedigreeFunction | Genetic influence |
| Age | Age |
| Outcome | 0 (No diabetes), 1 (Diabetes) |
⚙️ Procedure
- Load dataset
- Split into training and testing sets
- Train model without scaling
- Evaluate performance
- Apply feature scaling
- Train again
- Compare results
💻 Program (Scikit-learn Implementation)
Result
❌ Without Scaling
- Features with large values dominate (e.g., Glucose)
- Slower convergence of solver
- Suboptimal decision boundary
✅ With Scaling
- Balanced contribution of features
- Faster convergence
- Improved generalization
- Scikit-learn simplifies logistic regression implementation
-
Feature scaling significantly improves:
- Accuracy
- Precision & Recall
- Model stability
Comments
Post a Comment