founder & engineer · building leverage
writing

AI needs a playground where it can fail

Agents are now good enough to prove their own work end to end. What most teams are missing isn't a better model. It's the foundations: a safe, full-stack place where the agent is allowed to break things.

Since GPT-5.6, my development workflow has changed significantly. The surprising part is where that change came from. It didn’t come from the model writing better code, although it does. It came from the model getting good at proving that the code works.

That ability isn’t worth much on its own, though. A model can only prove its work if it has somewhere to run it: a real environment, with real services, real data, and room to break things. Most companies don’t have that. I think that’s the real bottleneck in AI-assisted engineering right now, and fixing it is foundational work, not model work.

The confidence gap

Before, my loop looked like most people’s. The agent implemented a feature and generated tests, and then I clicked through the app myself to confirm my expectations.

In a small codebase, that’s fine. In a complex application, it never gave me the confidence I needed.

The generated tests had a structural problem. The same model wrote the code and the tests, from the same understanding of the problem. If it misunderstood how invitations interact with permissions, the tests encoded the same misunderstanding and passed. The model was grading its own homework.

My manual click-through had the opposite problem. It touched the real system, but only the happy path, only once, and it left nothing behind. Two days later, I couldn’t tell you what I had actually checked.

Where confidence actually comes from.
approach real stack repeatable evidence
generated testschecks the code against itself
manual click-throughhappy path, once
agent-proven e2edrives the system, keeps receipts

What I was missing were real, fully fledged end-to-end tests: run against the actual stack, repeatable, and leaving evidence behind. I believe most companies aren’t prepared for them. They aren’t especially hard to write. The problem is that the environment to run them in doesn’t exist.

Let the agent prove it

So I changed the order. Now the agent runs end-to-end tests against our API before anything else happens, in two ways: automated, as an API consumer, and as a real user clicking through the app in a browser.

Along the way, it collects proof: database state, logs, metrics, UI screenshots. Only after I’ve reviewed that evidence does it generate the tests.

Prove first. Codify second.
  1. 01
    buildagent implements the change in its worktree
  2. 02
    runspins up its own full-stack environment
  3. 03
    provecalls the API as a consumer, clicks the UI as a user
  4. 04
    collectdb state · logs · metrics · screenshots
  5. gap found → back to 01
  6. 05
    codifyturn what was proven into tests

That last step matters more than it looks. Tests written after the proof aren’t guesses about how the system should behave. They record how it did behave, checked against the real thing. What it proved stays proven.

Evidence, not a green checkmark

The output of this loop isn’t a passing test run. It’s a receipt: a claim, plus the observations that back it up across every layer the feature touches.

An evidence receipt, illustrative.
claim "An invited teammate can accept and join the workspace."
apiPOST /invites → 201✓
dbinvite pending → accepted✓
uinew member in list, screenshot✓
logsinvite.accepted, 0 errors✓
metricsaccept latency within budget✓
apireplay accept → 200, expected 409!
5 proven 1 open question

This alone is gold. Reviewing an AI-generated change with a receipt like this is a completely different experience from reading a diff and hoping. I’m no longer asking “does this look right?” I’m asking “is this evidence sufficient?”, and that’s a question I can actually answer.

New models are insanely good at this. Give them the means to observe a system, and they go to considerable lengths to show you what happened instead of telling you.

The side effect: it questions the concept

Look at the last row of that receipt. To prove that an invite can be accepted, the agent had to think about what “accepting” means, including whether it should be possible to do it twice.

This is the side effect I didn’t expect. While agents test, they question the concepts around the feature. They find the race, the missing state transition, the permission that only makes sense on the happy path. These are issues you would usually never have found yourself, because you’re too close to your own mental model to see the gap.

Proving forces precision, and precision exposes every place where the spec was vague.

A playground where it can fail

Here’s the catch: none of this works without the right infrastructure.

An agent that proves its work has to run the whole system. It has to create users, send requests, break things, inspect the database, read the logs, and then do it all again. You can’t do that on a staging environment shared with everyone else, and you certainly can’t do it on production. A laptop with half the services mocked out only proves what the mocks allow.

What AI needs is a playground: a place where failing is safe.

At Leverage, we call these devloops. A devloop is an ephemeral, full-stack environment of our platform that spins up in one to two minutes. It comes with good defaults, test users, and skills: everything an agent needs to bring up its own environment per worktree, use it, and throw it away.

One worktree, one agent, one disposable stack.
worktree a feat/invites
webapidbqueue
seeded users · skills
worktree b fix/billing
webapidbqueue
broken db · no harm
worktree c infra/queue
webapidbqueue
seeded users · skills
ready in ~1–2 min isolated disposable

Isolation is what makes this scale. Three agents on three branches don’t fight over one database. A failed migration in one worktree is a non-event. The agent can be as destructive as the proof requires, because the environment is disposable by design.

The skills matter as much as the environment. An environment the agent doesn’t know how to drive is just expensive idle compute. The skills teach it how to seed data, where the logs live, how to sign in as a test user, and which metrics to look at.

This is the foundational work I mean. It isn’t glamorous: container images, seed scripts, service wiring, auth for test users, and documentation written for an agent instead of a human. But it’s the difference between an agent that claims its work is done and one that shows you.

Playgrounds for humans, too

The same idea extends beyond agents. On top of devloops, we have automated PR preview environments. Designers, and anyone else in the company, can try new things or test infrastructure changes end to end before anything lands on staging.

Our philosophy in a nutshell: test everything in the PR, merge to main, test on staging, then approve to prod.

Most of the proving happens before the merge.

Every stage has a job. The PR is where most of the proving happens, by agents and humans alike. Staging is a second look in a shared environment. Production gets a deliberate human decision. At our size, it works great.

Speed moves the bottleneck

Of course, speed brought its own problems.

With agents, we ship a lot. Keeping PRs up to date with main became really annoying, so we adopted GitHub merge queues. Then CI became the constraint, in both cost and time, so we optimized it for cheaper, faster cycles. Every fix moved the bottleneck somewhere else.

Solve one constraint, meet the next.
writing the codeagentssolved
trusting the codedevloops + proofsolved
stale PRsmerge queuessolved
CI cost & cycle timeoptimized CIsolved
feedback for the agentfast, deterministic signalsnow
??next

Now we’re looking for fast, deterministic quality signals to give the AI more feedback. End-to-end proof is the strongest signal we have, but also the slowest. The checks that tell an agent it’s wrong in seconds instead of minutes, and tell it the same thing every time, let it iterate many more times before it ever needs a devloop.

Foundations are the multiplier

Good engineering practices have never mattered more.

It’s tempting to think better models make engineering discipline less important. My experience is the opposite. The model is already capable of proving its work. Whether it actually can depends almost entirely on what you’ve built around it: reproducible environments, seeded data, observable systems, isolated previews, a clear path to production.

Those foundations are what turn a capable model into a trustworthy one. Give AI a playground where it can fail, and it will fail there, early and cheaply, instead of in front of your users.

AI should prove its work. Confidence comes from evidence.

I’m curious which limitation we’ll hit next.

references Leverage Computer GitHub docs: managing a merge queue