Replay
Full prior transcripts accumulate inside every later turn. Summing a single average turn misses that growth.
Compare full-history replay with a byte-stable cached prefix and bounded compact history. The curve shows what model choice cannot fix: tokens your code keeps sending again.
Compute the trace to reveal replay and cache-shaped growth.
Stable prefix + current user/retrieval + every prior transcript.
0 cached + 0 volatile
Quality and latency are not inferred; this tool prices only the supplied token trace.
Full prior transcripts accumulate inside every later turn. Summing a single average turn misses that growth.
A stable prefix earns cache-read pricing. Compaction caps the volatile history; retrieval and current user tokens remain full-price.
Replace the fixture with production traces. Cache eligibility, quality, latency, provider rules, and offload overhead still require live measurement.
116,800 baseline input tokens versus 24,000 cached + 34,800 volatile harness input tokens.
This is a disclosed synthetic trace, not a vendor benchmark. The result prices token assembly only and does not infer quality or latency.