Directional deployment planning

Fit the model before the benchmark.

Allocate hybrid precision across an MoE model, expose the binding constraint, and leave with a benchmark-ready engineering brief.

Calculating sample

Model and workload

Deployment inputs

Starting allocation

Current estimate

Fits with room to test

FIT
MXFP8
NVFP4
NF3
Total VRAM0 GBvs FP16
Headroom0 GBBinding constraint
Speed indexrelative planning estimate
Quality estimate0%floor 94%
Balanced tradeoff

Change an input to inspect the recommendation.

Scenario comparison

Three ways to deploy

Transparent assumptions

Know what is estimated

Weight residency

Uses nominal 8, 4, and 3-bit weights plus a 6% metadata allowance.

Runtime and KV cache

Approximates cache from active parameters, context, and concurrency; runtime adds 4.5% of weights plus 2.5 GB.

Quality and speed

Heuristic format penalties and multipliers support planning only. Validate on your own tasks and kernels.

Durable decision

Save the evidence

From estimate to evidence

Benchmark the constraint, not the whole unknown.

Your deployment brief will update with the current scenario.

Enjoy this tool? Build your own with Super