Capacity says it can load. Throughput says how fast.
Compare the analytical storage and matrix-math floors for a hypothetical streamed-weight design. Defaults are assumptions, not a COLIBRI or GLM benchmark.
Predicted bottleneck
Hypothetical source-scale example ready.
Analytical ceiling only. Real inference adds latency, cache misses, software overhead, attention, sampling, and other work. No speed, quality, cost, or hardware claim is verified here.
Active parameters74.40B
Streamed per token148.80GB
Compute work0.1488T
Token-rate ceiling≤0.0470/s
I/O BOUNDPrediction confirmed.
ScenarioReuseGB/tokenI/O secMath secCeiling
8× reuse: 18.60 GB/token · 2.657 s I/O · ≤0.3763 tok/s · I/O BOUND.
Why: 148.80 GB ÷ 7 GB/s = 21.257 s, versus 0.1488 TFLOP ÷ 15 TFLOP/s = 0.0099 s.
What it cannot say: this is not measured COLIBRI performance or proof about output quality, cost, or GPU comparison.