Prove the next layer.

Run words, meaning, and fused ranks against the same labels. Escalate only when the measured misses name a capability you do not have.

Labeled retrieval fixture

100%
AWAITING MODEL384D TARGET
SMALLEST PASSING SETUPNOT RUN

A recommendation appears only after every labeled query runs through all three methods.

LEXICALMRR —
SEMANTICMRR —
HYBRID RRFMRR —

Load the real local model and run the labeled sample.

Complexity needs evidence.

Exact-term misses justify lexical help. Paraphrase misses justify semantic help. Cross-document joins may justify graph or multi-step retrieval. None of those automatically justify an agent.

Fuse before you orchestrate

Reciprocal-rank fusion combines independent rankings without pretending their raw scores share one scale.

Label the real workload

A benchmark only guides architecture when queries, expected evidence, and document noise resemble production.

Escalate by missing capability

Images need multimodal parsing. Relations need graph evidence. Weak corpora need correction or abstention. Multi-step work needs decomposition proof.

COMPLETED RETRIEVAL BENCHMARK
MEASURED SETUP
No benchmark has completed.
Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.