Agentic systems / field notes

Loop Engineering: stop prompting the agent, build the system that prompts it.

A framework from a senior Anthropic engineer: autonomous agents that Discover their own work, Isolate it in branches, Execute fixes, and Verify before merging. No human in the prompt box.

DISCOVERISOLATEEXECUTEVERIFYMERGEthe loop

The Core Shift

Manual prompting scales with your attention. A loop system scales with your infrastructure.

before

Manual prompting

human → prompt → agent → paste result → repeat

  • You are the scheduler, router and reviewer
  • Work stops when you stop typing
  • Context is copied by hand, errors slip through
after

Loop system

signals → discover → isolate → execute → verify ↺

  • Failing CI and open issues are the prompt
  • Branches isolate every attempt
  • Tests gate every merge, humans handle escalations

Run the Loop

A mock repo has a failing CI run and three open issues. Press Play (or Step) to watch the agent discover work, branch, patch, verify and merge.

DISCOVERISOLATEEXECUTEVERIFYMERGE
acme/paymentsCI failing
  • #142 flaky test: retry logic in webhook handler
  • #141 upgrade stripe sdk to v14
  • #139 docs: rotate api keys guide
mainagent/fix-142
- retries = 0 # never retried + retries = 3 + backoff = exponential(base=0.5)
run log

idle — no run yet. Press Play to let the agent find its own work: it will read the failing CI signal, pick issue #142, branch, patch, verify and merge.

Design Principles

The loop only stays safe if these four properties hold.

Idempotent tasks

Re-running a task must be harmless. Loops retry; side effects must not compound.

Isolation

Every attempt lives on its own branch or sandbox. Failure is cheap and revertible.

Verification gates

Nothing merges without passing automated checks. The gate, not vibes, decides.

Human escalation

When the loop stalls or stakes are high, it files a report and pages a human.

Check Your Loop Literacy

Six quick questions. Instant feedback, score at the end.

Five scripted stages make a loop visible and reviewable

Read the explanation

This local demonstration has five actions: discover, isolate, execute, verify and merge. The discover text assigns three open issues and selects issue one hundred forty two. Execute reveals a mock diff with two added and one removed line. At eighty pixels per count five stages measure four hundred, two additions one hundred sixty and one removal eighty. These are browser elements and prewritten logs, not actual issue retrieval, branch creation or source patches. Verify sets the CI badge to passing and logs two hundred fourteen tests passed and zero failed. Merge shows a path and closes the illustrated issue, then logs the next candidate. Neither stage runs tests or invokes Git. At one pixel per test the assigned pass bar is two hundred fourteen and failure zero; at forty pixels per action all five stages measure two hundred. After the fifth stage both play and step buttons disable. Reset clears visible diff, merge and passing state so the same demonstration can be replayed. The quiz contains six questions and three options each, eighteen option buttons. Each question accepts one answer and increments score only for the coded correct choice. At fifteen pixels per option eighteen measures two hundred seventy while six questions at the same scale measure ninety. Score at least five selects the strongest feedback; three or four selects the middle message. This tests recall of the page concepts, not production agent reliability. Branch isolation, verification, idempotence and human escalation are design lessons; the scripted merge is not evidence that changes can safely ship without review.

Super generates helpful tools and automates fact-checking across the internet proactively. If you enjoyed this tool, build your own with Super and share it with a friend.