APPLIED ML ITERATION WORKBENCH

Tabular ML Iteration Lab

Phase 1 Validation Anchor StratifiedKFold
Fold 1
Fold 2
Fold 3
Fold 4
Fold 5
Optimal Split: Target distribution is strictly balanced across folds without identity spill.
Phase 2 30-Min Dumb Baseline
Baseline CV AUC 0.742 Benchmark Anchor
Train vs CV Gap 0.021 Healthy Margin
Practitioner Rule: Do not tune hyperparameters or add 50 features on day 1. Establish a fast local CV anchor first; every future engineering idea must beat this score.
Phase 3 Targeted Anomaly & Signal Inspector No Vanity EDA
Target Imbalance
9.4% Positives
PR-AUC / ROC
Missingness Signal
credit_score_null: 38%
Informative
Covariate Drift (p-val)
p = 0.384 (Stationary)
No Shift
Phase 4 Residual Error Analysis & Feature Crafting 3 Features Active
Filter for the highest validation false-positive and false-negative errors. Formulate a domain hypothesis, engineer targeted features, and inspect the immediate local CV gain:
ID Ground Truth Baseline Pred Current Pred Residual Error Primary Failure Cause
Phase 5 Late-Stage Model Blending Diversity Gain
Blend structurally distinct model predictions on the identical CV folds to cancel out uncorrelated residuals.
LightGBM (GBDT) 0.50
CatBoost (Oblivious Trees) 0.35
MLP (Neural Residuals) 0.15
Iteration Evaluation Low (Healthy Generalization)
Baseline CV AUC
0.742
Post-Feature CV
0.789
Ensembled CV AUC
0.798
Decile Err Reduction
18.4%
Train-CV Gap
0.021
Enjoy this tool? Build your own with Super