Model & Quantization
8B Q4_K_M
Model Parameters
8 Billion
Quantization Precision
FP16 (16-bit)
Q8_0 (8-bit)
Q4_K_M (4.5b)
Q2_K (2.5b)
Active Context Window
4,096 tokens
Batch Size / Streams
1 request
Hardware System Specs
Dedicated GPU VRAM
24 GB
System RAM (CPU Offload)
64 GB
Memory Bandwidth
800 GB/s
Economics & Usage
Local Hardware Investment
$1,600
Cloud AI API Subscription
$20 / mo
Daily Prompt Requests
500 req/day
Deployment Feasibility
100% VRAM
Fully fits in high-speed GPU VRAM
Gen Throughput Limit
132.8 t/s
Memory bandwidth bound
3-Year TCO Break-Even
N/A
Cloud API cheaper at current rate
Memory Allocation Breakdown
6.0 GB / 24.0 GB
VRAM Overhead vs Context Length
3-Year Cumulative Cost ($)
Deployment Launch Parameters
FROM llama3:8b-instruct-q4_K_M
# Hardware target: 24.0 GB VRAM | Context: 4096 tokens
PARAMETER num_ctx 4096
PARAMETER num_gpu 99