Hardware & Model Specs
Zero-API local math
VRAM Allocation Breakdown
6.53 GB / 24.00 GB
Model Weights (4.68 GB)
KV Cache (1.25 GB)
Runtime Overhead (0.60 GB)
Free Headroom (17.47 GB)
| Base Architecture & Parameter Count | 7.07 B (Qwen 2.5) |
| Attention Layers & GQA KV Heads | 28 Layers · 4 KV Heads (GQA) |
| Weight VRAM Memory Footprint | 4.68 GB |
| KV Cache Memory (Context: 8192, Batch: 1) | 1.25 GB |
| CUDA Context & PyTorch Runtime Overhead | 0.60 GB |
| Total Required Memory | 6.53 GB |
| Estimated Inference Throughput | 64.2 tokens/sec |
| Recommended Local Runner | llama.cpp / Ollama |
Agentic Paper Summarization Task Simulator
Demonstrates agentic multi-pass abstract evaluation using compact models (such as Inkling-Small or Qwen-2.5) on Hugging Face ICML 2026 paper reproductions.
Ready. Click 'Run Synthesizer' to simulate prompt ingest, KV cache allocation, and abstract synthesis.
Generated Deployment Script & Config
# Initializing deployment generator...