Every candidate must cover the same workload IDs. Unequal evidence is blocked before comparison.
Choose an agent pattern
from measured runs.
Compare identical workloads across source-named patterns. Reliability is a floor; latency or cost is the objective. No opaque score decides for you.
Benchmark run CSV
Minimize after the floor
Load the sample or paste run evidence. Analysis stays on this device.
One decision, inspectable evidence
Observed success stays beside its 95% Wilson interval, latency distribution, cost, and retries.
Clear the reliability floor first. Then minimize one named objective without a weighted score.
The source supplies a pattern catalog, not benchmark outcomes. This lens compares only the rows you provide and does not verify author experience, hidden reasoning, or live agent behavior.
Selected by evidence
No result yet
Re-run same rows:
Observed success / 95% Wilson--
Median / p90 latency--
Mean cost--
Mean retries--
Latency x cost tradeoff--
clears floorbelow floorcircle size is observed success
Pattern evidence
PatternSuccessMedianMean cost
Decision rule
Boundary
Export has not been created.