Interactive research explainer

Teaching language models to say “I’m not sure” — and mean it

Large language models often sound certain even when they are guessing. The paper concept Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs proposes rewarding models not just for being right, but for making their spoken confidence match their true internal confidence. This page lets you feel that idea by training a simulated model yourself.

MetacognitionCalibrationRL fine-tuningFaithful hedging
Open the training lab Take the hedge quiz

The faithfulness gap

A model has an internal confidence — a probability it assigns to its own answer. It also has an expressed confidence — the hedging language in its reply. When those two disagree,

Enjoy this tool? Build your own with Super