Can your judge agree with human gold?

Choose the stronger answer before seeing the label. Then trace groundedness, relevance, and simple heuristic evidence instead of trusting one vague quality score.

Refund boundary

GOLDEN CASE 1 / 3
User query

The rubric turns “good” into three scorable questions

GroundednessDoes every factual claim follow from retrieved context?
Answer relevanceDoes it directly resolve the user’s actual question?
Heuristic gateIs the answer direct, bounded, and free of unsupported promises?

Pick the better production answer

HUMAN AGREEMENT 0 / 0
Evidence traceChoose A or B to unlock

Commit before calibration.

A golden set is useful because its expected decisions stay hidden until evaluation.

?
This lab uses authored human labels and transparent rubric evidence. It does not run an LLM judge or embedding model; one small golden set cannot establish production reliability.
Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.