1. System Hardware & Stack
0 GB (CPU)48 GB
8 GB128 GB
264
2K32K
2. Hardware Allocation Matrix
OPTIMAL - Fits within 12GB VRAM with headroom
VRAM Footprint
5.80 / 12.0 GB
KV Cache Footprint
1.20 GB
System RAM Footprint
2.40 / 32.0 GB
Est. Tokens / Sec
42.5 t/s
Total Latency
1,470 ms
Node Inspector: Click any graph stage above
Latency: -- ms
Query Stage selected. Click "Simulate RAG" to watch vector search tokens & document chunk scoring pass through query expansion, SearXNG scraping, vector embedding, and context generation.
docker-compose.yml (Ollama + SearXNG + ChromaDB)
Generating code configuration...
Canonical Verification Proof Surface
LLM Model
Qwen2.5-7B-Instruct-Q4_K_M
VRAM Allocated
5.80 GB
KV Cache Allocated
1.20 GB
Tokens / Sec
42.5 t/s
Retrieval Latency
320 ms
Total Latency
1,470 ms
Feasibility Result
OPTIMAL - Fits within 12GB VRAM with 5.0GB headroom