LLM Spiky Generalization & RL Synthesizer

Verifiable RL Trajectories
Curriculum Presets:
Andrew Ho's Thesis: Standard LLMs suffer from "Spiky Generalization"—near-perfect accuracy on memorized/training patterns, but sharp performance drops on minor prompt variations. Synthesizing multi-step reinforcement learning rollouts with deterministic reward verification fills in the missing state space, producing smooth out-of-distribution performance curves.
1. RL Environment Tree Rollouts
Legend: Verified Path Pruned/Hallucination Active Trajectory
Select a node in the rollout tree to inspect state, action step, verifier status, and computed reward.
2. Generalization Performance Landscape
25%
State Space Coverage
0.88
Avg Verifier Reward
78%
OOD Generalization
12%
Pruned Noise Ratio
Enjoy this tool? Build your own with Super