| Model Tier | Assigned Tasks | Monthly Reqs | Token Footprint | Avg p50 Latency | Monthly Cost |
|---|
The default workload is five million monthly requests. Classification receives thirty percent, document extraction thirty five, security fifteen and reasoning twenty. Classification and extraction together form three million two hundred fifty thousand requests. Security has seven hundred fifty thousand and reasoning one million. At one pixel per ten thousand requests these bars measure three hundred twenty five, seventy five and one hundred. Mix sliders are normalized by their sum. These categories and model names are teaching fixtures, not verified live provider availability, pricing or capability. Default confidence zero point eight eight produces an escalation scale of zero point three eight divided by zero point four five. Classification escalates eight percent times that scale, while extraction escalates twelve percent times it. For five million requests this sends approximately one hundred one thousand three hundred thirty three classification requests and one hundred seventy seven thousand three hundred thirty three extraction requests onward, totaling two hundred seventy eight thousand six hundred sixty seven. At one pixel per thousand requests the bars match those counts divided by a thousand. This formula is an assigned rate; it never measures actual model confidence or evaluates individual task correctness. With dedicated security enabled, all seven hundred fifty thousand security requests go to that tier. Turning it off moves the same amount to the reasoning tier while leaving classification and extraction unchanged. At one pixel per ten thousand, dedicated security changes from seventy five to zero and added reasoning volume becomes seventy five. Cost is input and output tokens multiplied by assigned per-million prices, with fixed token reduction factors. Weighted p95 averages assigned tier p95 values; it is not the fleet ninety fifth percentile. Native exports serialize computed projections even if D3 visualization fails, so actual JSON export is verified separately from a healthy graph or real deployment.