| Paradigm Paper | Core Innovation | Harness Mechanism | Feedback / Reward Form | Primary Target Domain |
|---|---|---|---|---|
| Agent Lightning v1.0 | Lightweight, fast distributed RL harness orchestration | Multi-process RPC sandbox & mock broker | Sparse episode trajectory return + latency penalty | OS agents, web navigators, multi-tool search |
| LEGO-RL | Harness-native test composition for coding agents | Dynamic sub-unit test harness generator | Fine-grained execution pass/fail & assert density | SWE-Bench, Python refactoring, bug repair |
| EnvHarness | Awakens static documentation & APIs into interactive worlds | Synthetic schema-driven API simulator | State delta verification & schema consistency | Enterprise tool-use, database query reasoning |
| SPADE | Self-Play in Adaptive Synthetic Executable Arenas | Generator vs Solver adversarial curriculum | Elo differential & complexity-calibrated win reward | Mathematical logic puzzles, competitive planning |
For the coding fixture, the authored base pass value is sixty eight. Depth four adds eight point eight, while noise fifteen subtracts four point five, giving a pass-roll midpoint seventy two point three. Raising only depth to six gives seventy six point seven. At three pixels per midpoint percentage point the bars span two hundred sixteen point nine and two hundred thirty point one. A uniform random jitter from negative four to below four is then added and the result clamped from ten to ninety nine. No test cases, model outputs, environment transitions, or paper benchmarks measure these numbers. Names and research claims are source fixture descriptions, not independently verified results. The noise penalty multiplies the input divided by fifty by fifteen. Noise fifteen therefore costs four point five synthetic pass points, while noise fifty costs fifteen. At fifteen pixels per penalty point the bars span sixty seven point five and two hundred twenty five. Reward is a linear expression, two point three times pass fraction minus point eight. At midpoint seventy two point three it is point eight six two nine. These rules generate linked captions, not actual reward signals from action success or gradient learning. The source cycles five step indices and changes episode telemetry only when the index wraps to zero, rather than performing the stated policy and harness operations. The coding fixture subtracts loss decay point zero six times learning rate divided by point zero three five, then adds random jitter within point zero zero five. Starting loss point eight two at learning rate point zero three five has midpoint point seven six. Doubling learning rate to point zero seven gives midpoint point seven. At three hundred pixels per loss unit the bars span two hundred twenty eight and two hundred ten. A minimum loss point zero one five is applied. No model weights or gradient are updated. Native bindings occur before the missing D3 initialization failure; input/export can expose local state, while step highlighting errors can prevent displayed telemetry refresh. Separate exported synthetic values from a functioning graph, real reinforcement learning, exhaustive functions, or public proof.