LLM Quantization & Latency Simulator PTQ Engine

Model & Quant Config
VRAM Footprint
35.0 GB
↓ 75.0% Compression
Estimated Perplexity
5.42
+0.12 vs Baseline
Inference Speedup
3.1x
~112 tok/s (Memory Bound)
Weight Reconstruction MSE
0.00214
Quant Noise Residue
Weight Distribution & Quantization Grid
Pareto Efficiency: VRAM vs Perplexity
Calibration & Transformation Formulae
Scale Factor (S): S = (x_max - x_min) / (2^b - 1) = 0.023529
Zero Point (Z): Z = round(-x_min / S) = 0
Quantization Function: q = clamp(round(w / S) + Z, q_min, q_max)
Dequantized Weight Approximation: w' = S * (q - Z)
Simulation Proof Status
Initialized: 70B Model | INT4 Target | Group-128 Scaling | VRAM: 35.0 GB | PPL: 5.42 | MSE: 0.00214
Enjoy this tool? Build your own with Super