Test what your retriever remembers.

Measure source recall and the courage to stop. No answer generation, no hidden reranking, no invented confidence.

Evaluation inputs

0.22

Semantic evidence ledger

MODEL NOT READY
-ANSWERABLE RECALL
-UNKNOWN GATE
-CASES PASSED
-CONTEXT REDUCTION
Run the labeled benchmark to create real vectors and source rankings.
Waiting for local model

Retrieval is a testable subsystem.

Recall asks whether known evidence lands on the right source. Abstention asks whether missing evidence stops the pipeline. Both must be measured before fluent generation can hide the failure.

Similarity is not truth

Cosine scores rank candidates. They do not verify the facts inside a passage.

Thresholds are local

The bundled gate is a test fixture, not a universal constant. Calibrate it on representative positives and negatives.

Evaluation precedes scale

A small labeled set reveals wrong-source routing before thousands of files make failures harder to inspect.

COMPLETED RETRIEVAL AUDIT

Audit pending

-source recall
-unknown gate
-cases passed

This audit does not generate answers. Similarity is not truth, thresholds require corpus-specific calibration, and four examples do not establish production quality.

Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.