The faithfulness gap
A model has an internal confidence — a probability it assigns to its own answer. It also has an expressed confidence — the hedging language in its reply. When those two disagree,
Interactive research explainer
Large language models often sound certain even when they are guessing. The paper concept Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs proposes rewarding models not just for being right, but for making their spoken confidence match their true internal confidence. This page lets you feel that idea by training a simulated model yourself.
A model has an internal confidence — a probability it assigns to its own answer. It also has an expressed confidence — the hedging language in its reply. When those two disagree,