Choose a measured operating point for the exchange.

Parse labeled probe and classifier traces, quantify selective escalation, and defend a two-stage monitoring policy with observed recall, precision, and compute.

Paper reference, not reproductionThe source reports about 5.5% escalation, a 0.55/0.45 blend, and roughly 3.5% relative compute in one deployment. This workbench evaluates your data and does not treat those values as universal targets.

Evaluation traces

Required columns: id, context, probe_score, classifier_score, label, token_index.

Load the labeled sample or import a CSV.

Routing threshold sweep

Each point recomputes the cascade on the parsed labels. The vertical line is the active route threshold.

0.62
0.60
0.55
18
RecallPrecisionEscalation
No evaluated policy yet.

Observed policy result

RecallTP / positives
PrecisionTP / flags
Escalationrouted / rows
Relative computevs classify every row
False negativespositive rows missed
Routing ruleProbe score must meet the route threshold before classifier cost is spent.
Decision ruleRouted rows blend both signals; unrouted rows retain the probe score.

Constraint selector

Find the highest route threshold that satisfies both measured constraints.

Evaluate labeled traces before selecting a policy.

The aggregate only earns trust when every row is inspectable.

Review which exchanges paid for the second stage, how signals blended, and where labels disagree with the final decision.

IDContextProbeClassifierLabelRoutedFinal scoreDecisionOutcome
No evaluated rows.

Read the policy without turning evidence into a guarantee.

The workbench computes validation-set behavior. Deployment drift, adaptive attacks, and unlabeled edge cases remain outside this result.

More rows purchase the second-stage classifier. That can recover positives whose probe scores sit below a stricter threshold, but it increases compute and can still introduce false positives depending on the blend.

Preserve the measured policy.

The export contains settings, assumptions, aggregate metrics, the threshold sweep, and row-level decisions.

Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.