Polynomial Curve Fitting Playground

Why ML models underperform: too little flexibility misses the signal (underfitting), too much memorizes the noise (overfitting). Explore both live.

Fit It Yourself

Train MSE -Holdout MSE -

Click empty space to add a point. Drag points to move them. Hidden truth is a cubic; blue points are train, hollow points are holdout.

The Bias-Variance U-Curve

Train error (blue) falls monotonically with degree. Holdout error (orange) traces a U: high in the underfit zone, lowest at the sweet spot, rising again as overfitting sets in. The dot marks your current degree.

Worked Example: Degrees 1, 3, 12

Degree 1 — Underfit

A line cannot bend with the cubic signal. High bias: both train and holdout error stay large.

Degree 3 — Sweet spot

Matches the true generating process. Low bias, low variance: holdout error near its minimum.

Degree 12 — Overfit

Wiggles through every training point, chasing noise. Train error near zero, holdout error explodes.

Check Your Understanding

Overfitting FAQ

What is overfitting?

A model that fits training data too closely, capturing random noise instead of the underlying pattern, so it generalizes poorly to new data.

How do I detect it?

Compare train and holdout error. A large gap, with train error far lower, signals overfitting.

How do I fix underfitting?

Increase model capacity: higher polynomial degree, more features, deeper networks, or less regularization.

Why does holdout error form a U shape?

Error decomposes into bias plus variance. Bias falls as capacity grows while variance rises; their sum is minimized at an intermediate complexity.

How polynomial capacity and noisy observations change generalization

Read the explanation

The source creates twelve noisy training points and eight independent holdout points around the cubic shown here. This video uses one particular source-generated draw, whose seed is recorded in the private evidence. Blue is training error and yellow is holdout error. Both share the original chart’s logarithmic error scale. The fit uses the source normal equations with a tiny diagonal regularizer, not an ideal unregularized solver. Move through the actual integer degree settings, from one through twelve. The cursor joins adjacent evaluated degree results; a fractional cursor position does not create a fractional polynomial degree. Watch training error fall while holdout error changes separately. The degree-three example reduces both errors compared with the line. Increasing capacity further nearly removes training error, but does not remove holdout error. In this draw, degrees one, three and twelve have holdout errors about point nine eight five, point five five nine and point nine two one. Twelve therefore generalizes worse than three here, even though its training error is much smaller. This is a worked finite example. Other noise draws can give different shapes; the sidebar’s sweet-spot label alone does not prove an optimum. Now keep the twelve original training points and add one more point at horizontal coordinate zero. Fix the degree at one, so the model is a line. The added point’s vertical coordinate is the input. The yellow curve is the added value and the blue curve is the fitted line’s prediction at zero, which is its intercept. Both use the same vertical coordinate units. Drag that added training point from minus one to two. Recompute the original source solver at every input value. The intercept responds, but it does not follow the added value one for one: all thirteen training points enter the same normal equations. The two markers share the input axis, so their different movements show how a single observation influences the whole fitted model. The fitted response here is affine because the training horizontal coordinates stay fixed while one target changes. This is a source-model sensitivity experiment, not a claim that the added point is correct data. In the real tool you can add and drag points yourself and compare the equation and error readouts. The tiny diagonal regularizer remains part of this calculation. For the last comparison, return to the original twelve points and set degree twelve. Move only the seventh training point, at horizontal coordinate about point one two three’s vertical coordinate; keep every horizontal coordinate and all eight holdout points fixed. The input axis is that point’s value. Blue and yellow again mean training and holdout error, on a shared logarithmic error scale. Drag the seventh value through a range of plus or minus point eight around its original value. The source solver can nearly match the training targets throughout this interval. Yet the holdout error changes as the coefficients respond to that one training point. These curves are recomputed from actual source formulas; a low training residual does not certify a stable prediction on unseen points. The holdout display follows the original chart’s cap at fifty, while the numeric evidence retains uncapped errors. This example isolates sensitivity at one fixed capacity, not a universal risk forecast or proof of convergence. The native tool also lets you randomize noise, reset the point sets and answer five feedback questions. The original controls and source randomness remain unchanged.

Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.