Binary working-set optimizer
Put heat where memory is.
Profile routed experts, set a GPU memory cap, and solve the integral placement that maximizes activation-weighted latency savings.
Use measured data in production. Shares must sum to 1.00. This model excludes transfer, contention, batching, NUMA, cache, and multi-expert overlap.
Load the bundled synthetic profile or paste measured JSON.
—GPU memory used
—activation coverage
—expected saving
—CPU activation tail
GPU working set
CPU capacity tail
Integral optimum
—Maximize summed share × (CPU ms − GPU ms), subject to footprint ≤ budget.Boundary
Synthetic until replacedThe sample proves the optimizer lifecycle, not KTransformers or hardware benchmark performance.