Agent Engineering

The Agent Evaluation Flywheel

Frameworks like AgentLoop close the loop between running agents and improving them: observe trajectories, turn traces into datasets, judge with an agent, and feed the lessons back into memory. Spin it faster and quality compounds.

FLYWHEEL RPM · 0 cycles/week · drag to orbit
STATION 01

Trajectory Observability

    Automation level drives the flywheel

    Manual Automated 20%
    StageHours / cycle
    1. Observe trajectories-
    2. Trace to dataset-
    3. Judge results-
    4. Update memory-
    Total per cycle-

    Iteration throughput

    -
    improvement cycles per 40h work-week
    Cycle time-
    Cycles per quarter-
    Speedup vs fully manual-
    Human eval hours saved / week-

    Model: each stage has a manual cost and an automated cost; the slider linearly interpolates. Cycles/week = 40h divided by total cycle hours. Illustrative numbers, but the shape is real: judging is the bottleneck, so automating it moves the needle most.

    Agent-as-a-Judge vs human evaluation

    Human review is the gold standard for judgment quality but it does not scale: a careful reviewer covers 20-50 trajectories a day, gets fatigued, and drifts. An LLM judge scores thousands per hour at near-zero marginal cost, which is what makes a flywheel possible at all.

    Consistency metrics that matter

    Before an automated judge replaces human eval, teams verify it is consistent with humans and with itself.

    The four stations, in practice

    Why flywheels beat one-off evals

    A static benchmark measures an agent once; a flywheel improves it continuously. The compounding comes from the loop, not any single stage.

    Worked example: what one extra cycle per week is worth

    Suppose each improvement cycle fixes failure modes worth a 2% absolute gain in task success rate, with diminishing returns. At 1 cycle/week (fully manual judging) an agent at 70% success reaches roughly 78% in a month. At 8 cycles/week (automated judging with human audits) the same team runs 32 cycles and plateaus near its data-quality ceiling in the same month, then spends the rest of the quarter expanding coverage instead of waiting on reviews.

    The arithmetic behind the slider above:

    Two honest caveats. First, automation quality gates the whole loop: a judge with kappa 0.5 against humans produces fast noise, not fast learning. Second, the flywheel measures what the judge can see; silent failure modes (subtle factual errors, slow degradation of tone) still need scheduled human deep-dives. Mature teams typically settle around 80-90% automation with a standing 5-10% human audit sample.

    Enjoy this tool? Build your own with Super