1. LLM Training Pipeline Simulation
Simulate pre-training ingest, supervised tuning (SFT), and RLHF alignment.
Self-Attention Token Salience
Softmax(QK^T / sqrt(d_k))
Cross-Token Attention Matrix Heatmap
Query vs Key Weights
Attention Entropy
0.142
Perplexity (PPL)
8.24
Curriculum Step
12,450 / 25k
2. Iterative Prompt Engineering Sandbox
Inject context, enforce negative constraints, and score model alignment.
Aligned Production Prompt
94.5%
Compose a revised letter to Robert addressing vehicle liability and car damage claims. Specifically, remove Section II while strictly preserving all other sections and original tone.
Alignment Evaluation Breakdown
RLHF Reward Model Weights
Context Grounding & Specificity:
96.0%
Negative Constraint Retention (Section II guard):
98.2%
Hallucination / Entropy Penalty:
-3.7%