Adaptive thinking replaces extended thinking — on by default in Claude Code and the API. The model decides when to think; you set how hard. Drag the reasoning core to orbit; change effort to reshape it.
Before answering, the model can emit an internal chain of thinking tokens — scratch-work it uses to plan, check edge cases, and self-correct. You pay for them like output tokens, and they add latency before the visible answer starts.
More thinking helps on genuinely hard steps (multi-file reasoning, tricky math, long plans). It adds nothing on trivial ones.
Extended thinking was a switch plus a fixed token budget: you guessed up front how much scratch-work every request needed.
Adaptive thinking is on by default in Claude Code and the API. The model decides per-step whether to think at all; your effort setting (low / medium / high) only biases how deep it goes when it does. No budget guessing.
| Task | Effort |
|---|---|
| Quick edits, renames, lookups | low / medium |
| Most day-to-day coding | medium (default) |
| Long agentic runs, many tools | high |
| Memory-heavy, cross-session work | high |
Start at medium. Bump to high only when runs are long or state is heavy — most coding won't need more.
Every thinking token is billed and delays the answer. Forcing high effort on a one-line fix can multiply cost and latency for identical output quality.
Rough intuition: cost ≈ input + thinking + output tokens. Only one of those is under a dial — spend it where reasoning depth actually changes the result.
In the API you pass an effort hint alongside your request, e.g. thinking: { effort: "high" } — no more token budgets to size. In Claude Code, adaptive thinking is simply on; the default effort is medium and you only override it for special workloads.
Because the model skips thinking on easy steps, a whole session at high is far cheaper than the old "max budget on every call" pattern — depth is spent only where the model judges it's needed.
You're kicking off an overnight autonomous agent that will chain hundreds of tool calls across a large repo. Which effort do you pick?
This page illustrates effort by swapping fixed visual settings. Low displays six hundred particles and one ring. Medium displays fifteen hundred particles and three rings. High displays twenty six hundred particles and five rings. At point one five pixels per particle, the medium and high bars grow to two hundred twenty five and three hundred ninety pixels. Nothing here counts actual model reasoning tokens. These saved constants produce an animation, not a measured account of a model internal process or a claim about current product defaults. The cost label changes from one point zero times at medium to two point five times plus at high. The gauge width separately changes from thirty eight to eighty eight percent. Their ratio is approximately two point three two, not two point five. The two bars shown here reproduce the saved gauge percentages on a four hundred pixel track, reaching one hundred fifty two and three hundred fifty two pixels. This distinction matters because the gauge is an authored display choice. It is not a billing formula, a price forecast, or verified current provider behavior. Clicking a task updates recommendation text and pulses the suggested effort button. It does not itself call the effort switching function, so a high recommendation can appear while the actual dial remains medium. Clicking the effort button commits the visual change. A correct quiz answer separately does call high effort. The two bars contrast the saved medium and high active ring counts, three and five, at sixty pixels per ring. In an offline browser without the Three library the renderer fails before those handlers bind, so the local checks can establish preserved failure behavior rather than a healthy three dimensional interaction.