2D Prompt Embedding Space & Cosine Halo Radius
Normalized Embedding Projection (PCA/t-SNE) Cached Seed Vector
Semantic Cache Hit
Cache Miss (Calls LLM)
Cosine Similarity Radius
Execution Replay & Dynamic Decision Table
7 Queries Loaded| Prompt Text | Cluster | Exact Redis | Semantic Sim | Max Cosine | Latency |
|---|
Standard vs Semantic Cache
Standard Exact Match
0 / 7
Hit Rate: 0.0%
Cost Saved: $0.000
Semantic Vector Cache
4 / 7
Hit Rate: 57.14%
Cost Saved: $0.0180
Latency Benchmark: 1,800ms (LLM OpenAI call) vs 38ms (Semantic Cache return).
Cost Benchmark: $0.0045 average cost per query (GPT-4o / Claude 3.5 Sonnet token burn).
Production Code & Config
Numpy vector cosine math ready for Redis or Milvus.