Interactive fit
Twelve blue points were sampled from a hidden cubic y = a0 + a1x + a2x2 + a3x3 plus noise. Click empty space to add a point, drag any point to move it, then change the polynomial degree to refit by least squares.
Add at least 2 points by clicking the chart to fit a curve.
Train vs holdout error
Training error keeps falling as degree grows, but error on held-out points traces the classic U shape: high in the underfit zone, lowest near the true degree, and rising again as the model chases noise.
Worked example: degree 1 vs 3 vs 12
Degree 1 underfit
A straight line cannot bend with the cubic. High bias: it is wrong on train and holdout alike.
Degree 3 sweet spot
Matches the true generating process. Low bias, controlled variance, best holdout error.
Degree 12 overfit
Wiggles through every noisy point. Tiny train error, large holdout error: high variance.
Check your understanding
FAQ on overfitting
Why do ML models underperform in production?
Often because they overfit training data: they learned noise specific to the training set, so accuracy collapses on new data. Comparing train and holdout error, as this page does, is the standard diagnostic.
What is the bias-variance tradeoff?
Bias is error from an overly simple model; variance is error from sensitivity to noise. Increasing degree lowers bias but raises variance. The best model balances the two at the bottom of the U-shaped holdout curve.
How do I pick the right polynomial degree?
Hold out data (or cross-validate), fit each candidate degree, and choose the one with lowest holdout error rather than lowest training error.