Trace the memory movement active parameters leave out.

Model the resident set, replay routed experts, and compare reactive LRU against an oracle lookahead bound. Every assumption stays visible.

CONSUMER HARDWARE MODELTRANSFER-ONLY ESTIMATE · NOT A LIVE BENCHMARK
Pairs separated by |
Expert capacity4 slots6.0 GB available after shared + KV
LRU transfers18.0 GB12 misses · 4 hits
Lookahead transfers15.0 GB10 misses · 6 hits
Modeled stall reduction16.7%1500 ms → 1250 ms

Same 16 activations. Different movement.

REACTIVE LRU

12 misses

Evicts the least recently used expert with no knowledge of upcoming routes.

ORACLE LOOKAHEAD

10 misses

Evicts the resident expert needed farthest in the future: an upper bound, not a predictor claim.

Each bar is transferred expert weight per token. Coral is LRU; green is lookahead.

Interpretation. Active weights describe one instant, not the path. Shared layers and KV state consume the fixed budget; route transitions determine misses; bandwidth converts movement into a transfer-only stall estimate. Compute overlap, page faults, quantization metadata, attention compute, and predictor error remain outside this model.

RESULT READY · 16 ACTIVATIONS SIMULATED

The populated export preserves assumptions, both policy summaries, and every hit/miss event.

Lookahead removes 16.7% of modeled transfer stall.

4EXPERT SLOTS
18 → 15 GBLRU TO LOOKAHEAD MOVEMENT
1500 → 1250 msTRANSFER-ONLY STALL

The activation count stays at 16. The improvement comes from retaining experts whose next use is closer, not from pretending inactive weights disappeared. Real predictive prefetch will sit below this oracle bound and must be benchmarked on the target machine.

Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.