Fuse before you orchestrate
Reciprocal-rank fusion combines independent rankings without pretending their raw scores share one scale.
Run words, meaning, and fused ranks against the same labels. Escalate only when the measured misses name a capability you do not have.
A recommendation appears only after every labeled query runs through all three methods.
Load the real local model and run the labeled sample.
Reciprocal-rank fusion combines independent rankings without pretending their raw scores share one scale.
A benchmark only guides architecture when queries, expected evidence, and document noise resemble production.
Images need multimodal parsing. Relations need graph evidence. Weak corpora need correction or abstention. Multi-step work needs decomposition proof.