The Gradient Descent Math Lab

Knowing that gradient descent works gets you an import statement. Knowing why it works — curvature, conditioning, step sizes — is what makes you stand out in AI/ML. Roll a ball down three real loss surfaces and watch the math misbehave.

step 0
θ = (0.00, 0.00)
L(θ) = 0.000
‖∇L‖ = 0.000
converging

The Update Rule

v ← β·v − η·∇L(θ)    θ ← θ + v η=0.020 β=0.00 → Δθ=(0,0)

Surface: Convex Bowl

Regimes to Recognize

Convergeη < 2/λmax (largest curvature): steps shrink, L → minimum smoothly.
Oscillateη near the stability limit: the ball ping-pongs across steep walls while creeping along shallow ones — the classic ill-conditioning picture.
Divergeη > 2/λmax: each step overshoots by more than it corrects; L blows up. On a quadratic this bound is exact — provable, not folklore.

Why the math matters

Ball = parameters θ; the arrow is −∇L, always the locally steepest descent direction (first-order Taylor: L(θ+d) ≈ L(θ)+∇L·d).
Trail = optimization trajectory. Momentum (β) accumulates past gradients — a leaky integrator that damps oscillation and accelerates along consistent directions.
Saddle points, not local minima, dominate high-dimensional losses; gradient noise (SGD) helps escape them. That result comes from random-matrix theory — more evidence the theory pays rent.

A 3-line convergence proof sketch

For L(θ) = ½θᵀAθ (A symmetric positive-definite), one GD step gives θ⁺ = (I − ηA)θ. Errors shrink iff every eigenvalue of (I − ηA) has magnitude < 1, i.e. |1 − ηλᵢ| < 1 for all i, which rearranges to 0 < η < 2/λmax. The slowest mode contracts by factor (1 − ηλmin) — so the iteration count scales with the condition number κ = λmax/λmin. Try the Bowl surface (κ = 4): the shallow x-axis is always the last to settle. This one derivation explains learning-rate warmup, why normalization layers help (they improve conditioning), and why Adam rescales per-coordinate.

Batch GD vs SGD vs Mini-batch

Batch GD (shown here) uses the exact gradient — deterministic, but each step costs a full pass over the data. SGD estimates ∇L from one sample: unbiased (E[ĝ] = ∇L) but noisy; the noise acts like temperature, helping escape saddles and sharp minima. Mini-batch (the practical default, 32–512 samples) trades variance ∝ 1/batch for hardware parallelism. Classic Robbins–Monro conditions for SGD convergence: Σηt = ∞ and Σηt² < ∞ — the reason learning-rate schedules decay over time rather than staying constant.

Read curvature before choosing the descent step

Read the explanation

The source bowl is one half x squared plus two y squared. Its gradient is x and four y, so its Hessian eigenvalues are one and four. With zero momentum, a descent step multiplies coordinates by one minus learning rate times curvature. At learning rate zero point two, x contracts by zero point eight and y by zero point two. The higher curvature direction responds more strongly. For this quadratic with zero momentum, strict contraction requires absolute one minus learning rate times each eigenvalue to be less than one. The largest eigenvalue four gives a learning rate below one half. At one half, the y factor is minus one, alternating without contraction; at zero point six it is minus one point four and expands. This threshold describes this particular quadratic without momentum, not every optimizer or loss. The source saddle is one half x squared minus one half y squared. At zero zero the gradient is zero, but opposite curvatures make it a saddle. At point one one and learning rate zero point two with no momentum, the next point is zero point eight one point two. Loss changes from zero to minus zero point four. The moving point shows descent escaping in the negative-curvature direction; lower loss does not prove a minimum.

Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.