Errors can go quiet without going away.
Route the same 60 claims through evidence and abstention. Watch confidence split from correctness, then see which errors a fluent interface leaves unnoticed.
Confidence is a display. Calibration is a measurement.
A system is calibrated when claims shown at 80% confidence are correct about 80% of the time. Verification can detect unsupported claims before display. Abstention can refuse the rest. Both reduce visible errors, but neither turns confidence itself into evidence.
That is the missing variable in the original question. Fewer complaints can mean fewer wrong answers, better filtering, more abstention, lower notice, or a mixture. This synthetic lab does not claim which one explains the whole industry; it teaches how to tell the mechanisms apart.
Transfer the gate
An AI gives a confident health claim, but two reputable sources conflict. What policy carries the lesson into this higher-stakes setting?
Synthetic educational model, not a benchmark and not a live hallucination rate. Verification-detection percentages are explicit teaching assumptions. High-stakes factual decisions still require qualified sources and domain professionals.