Honest Fine Print
Hardware reality for a 27B coder
A 27B-parameter model at 4-bit quantization needs roughly 16–20 GB of VRAM (approximate) — a 24 GB card (e.g. used 3090-class) runs it; at 8-bit or long contexts you need more, or you spill to slow CPU offload.
When the API still wins
Bursty/low usage (break-even never arrives), needing frontier-level quality — local 27B models are strong but measurably behind top hosted coders on hard tasks — zero maintenance, instant scaling, and no upfront cash. Your time debugging drivers is a real cost this lab doesn't price.
When local wins
Sustained heavy token volume, privacy/air-gap requirements, latency-insensitive batch work, and hardware you already own (sunk cost → only electricity matters).