Python Semantic Cache Workbench Cosine Proximity Sim
LLM Dollars Saved $0.0180
Latency Saved 7.05s
Semantic Savings % 57.14%

2D Prompt Embedding Space & Cosine Halo Radius

Normalized Embedding Projection (PCA/t-SNE)
Cached Seed Vector
Semantic Cache Hit
Cache Miss (Calls LLM)
Cosine Similarity Radius
Cosine Similarity Threshold (τ): 0.82 Default recommended: 0.80 - 0.85

Execution Replay & Dynamic Decision Table

7 Queries Loaded
Prompt Text Cluster Exact Redis Semantic Sim Max Cosine Latency

Standard vs Semantic Cache

Standard Exact Match
0 / 7
Hit Rate: 0.0%
Cost Saved: $0.000
Semantic Vector Cache
4 / 7
Hit Rate: 57.14%
Cost Saved: $0.0180
Latency Benchmark: 1,800ms (LLM OpenAI call) vs 38ms (Semantic Cache return).
Cost Benchmark: $0.0045 average cost per query (GPT-4o / Claude 3.5 Sonnet token burn).

Production Code & Config


          
Numpy vector cosine math ready for Redis or Milvus.
Enjoy this tool? Build your own with Super