Why the denominator shrinks
A length-stopped answer is not clean evidence of model quality. The gate preserves it, names the item, and removes the entire pair from both accuracy denominators so baseline and candidate remain comparable.
Normalize mixed answers, withhold truncation, and compare only paired evidence. The bundled fixture is synthetic and local; no model is trained or queried.
Protocol fields: run, item_id, task, answer_format, expected, answer, finish_reason, max_tokens. Lists are order-independent; object keys are canonicalized.
| Item | Format | Baseline | Candidate | Evidence decision |
|---|
—
A length-stopped answer is not clean evidence of model quality. The gate preserves it, names the item, and removes the entire pair from both accuracy denominators so baseline and candidate remain comparable.
Reporting 75% candidate accuracy across all four items would mix three completed answers with one truncated response. The paired gate reports 100.0% on three comparable items and keeps the +66.7 pp gain provisional until q3 is rerun.