OxAlpha Agent Latency & Throughput Simulator Z.ai GLM Multi-Step Inference

Effective Speedup
3.77x
OxAlpha Runtime
39.4s
Standard Runtime
148.6s
Prefill Latency Saved
64.2s
Peak VRAM Saved
14.8 GB
Step 1 / 12: AST Parsing & Symbol Resolution
Cumulative Context: 2,350 tokens
OxAlpha Step Latency: 3.12s (Std: 11.8s)

Multi-Step Execution Pipeline Log

Throughput from fixed latency assumptions

Read the explanation

The default twelve-step scenario generates four hundred fifty plus three times one hundred eighty tokens per step, totaling nine hundred ninety. It assumes standard generation at thirty eight tokens per second, fast generation at one hundred forty two multiplied by a speculative boost of one point two five, and fixed dispatch delays. Summing the source formulas gives about ninety four point five seconds fast and six hundred thirteen point five standard. At point three five pixels per estimated second the bars span thirty three point zero seven five and two hundred fourteen point seven two five. These are arithmetic assumptions, not benchmark measurements or actual model calls. With prefix caching enabled, every fast prefill uses eighteen hundred fifty divided by twelve hundred plus point one, about one point six four one seven seconds, even for the first step. Turning caching off uses the full prompt count of five hundred plus previous steps times eighteen hundred fifty. Across twelve steps, total estimated fast latency rises from about ninety four point five to one hundred eighty one point six seconds. At one point two pixels per second the bars span one hundred thirteen point four and two hundred seventeen point nine two. This fixed cached formula is a source simplification, not proof that a real runtime can reuse every prefix or that these timings generalize. The speculative checkbox multiplies the entered fast generation rate of one hundred forty two by one point two five, giving one hundred seventy seven point five effective tokens per second. Turning it off restores one hundred forty two. At one point three pixels per assumed token per second the bars span two hundred thirty point seven five and one hundred eighty four point six. The gain is hard-coded arithmetic rather than measured speculation acceptance. KV compression separately scales modeled retained prompt tokens, leaving generated tokens uncompressed and using a supplied memory multiplier. The JSON export captures this local estimated trace, not live hardware telemetry, a verified speedup, or a production model evaluation.

Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.