Knowing that gradient descent works gets you an import statement. Knowing why it works — curvature, conditioning, step sizes — is what makes you stand out in AI/ML. Roll a ball down three real loss surfaces and watch the math misbehave.
| Converge | η < 2/λmax (largest curvature): steps shrink, L → minimum smoothly. |
| Oscillate | η near the stability limit: the ball ping-pongs across steep walls while creeping along shallow ones — the classic ill-conditioning picture. |
| Diverge | η > 2/λmax: each step overshoots by more than it corrects; L blows up. On a quadratic this bound is exact — provable, not folklore. |
For L(θ) = ½θᵀAθ (A symmetric positive-definite), one GD step gives θ⁺ = (I − ηA)θ. Errors shrink iff every eigenvalue of (I − ηA) has magnitude < 1, i.e. |1 − ηλᵢ| < 1 for all i, which rearranges to 0 < η < 2/λmax. The slowest mode contracts by factor (1 − ηλmin) — so the iteration count scales with the condition number κ = λmax/λmin. Try the Bowl surface (κ = 4): the shallow x-axis is always the last to settle. This one derivation explains learning-rate warmup, why normalization layers help (they improve conditioning), and why Adam rescales per-coordinate.
Batch GD (shown here) uses the exact gradient — deterministic, but each step costs a full pass over the data. SGD estimates ∇L from one sample: unbiased (E[ĝ] = ∇L) but noisy; the noise acts like temperature, helping escape saddles and sharp minima. Mini-batch (the practical default, 32–512 samples) trades variance ∝ 1/batch for hardware parallelism. Classic Robbins–Monro conditions for SGD convergence: Σηt = ∞ and Σηt² < ∞ — the reason learning-rate schedules decay over time rather than staying constant.
The source bowl is one half x squared plus two y squared. Its gradient is x and four y, so its Hessian eigenvalues are one and four. With zero momentum, a descent step multiplies coordinates by one minus learning rate times curvature. At learning rate zero point two, x contracts by zero point eight and y by zero point two. The higher curvature direction responds more strongly. For this quadratic with zero momentum, strict contraction requires absolute one minus learning rate times each eigenvalue to be less than one. The largest eigenvalue four gives a learning rate below one half. At one half, the y factor is minus one, alternating without contraction; at zero point six it is minus one point four and expands. This threshold describes this particular quadratic without momentum, not every optimizer or loss. The source saddle is one half x squared minus one half y squared. At zero zero the gradient is zero, but opposite curvatures make it a saddle. At point one one and learning rate zero point two with no momentum, the next point is zero point eight one point two. Loss changes from zero to minus zero point four. The moving point shows descent escaping in the negative-curvature direction; lower loss does not prove a minimum.