Artificial network
A value estimate changes in proportion to a prediction error. Every number remains visible and inspectable.
Change a reward expectation, expose the prediction error, and watch one idea travel between reinforcement learning and dopamine research.
The artificial lane exposes exact temporal-difference arithmetic. The biological lane offers a careful interpretation: changes in dopamine activity can carry information about outcomes relative to expectations, not simply pleasure.
A value estimate changes in proportion to a prediction error. Every number remains visible and inspectable.
Phasic dopamine responses are widely studied as a candidate teaching signal. The mapping is useful, but brains are richer than this equation.
A burst can reflect better-than-expected outcomes; a dip can reflect worse-than-expected ones. Expectation is the hinge.
Dopamine is not a happiness meter. In this lab, it is a lens on surprise: the gap between what a learning system expected and what arrived.
In the artificial lane, the positive error strengthens the value estimate. In the biological lens, a phasic dopamine burst is often interpreted as a better-than-expected teaching signal.
When an outcome is better than expected, the positive error increases the learned value. Repeated trials shrink the error as the expectation catches up.
If a reward was expected but does not arrive, the error turns negative. The model revises its estimate downward; the biological interpretation points to a dip from baseline.
The equation is a compact learning rule. The brain includes multiple circuits, timescales, and functions. Use the comparison to ask sharper questions, not to flatten biology into code.