← Back to Notes
· 5 min read / AI, Risk, Engineering

Execution Is Free Now. Verification Isn't.

The cost of making things is collapsing. The cost of trusting them isn't. The widening gap between what a machine can produce and what a human can actually verify is where the next decade of leverage, and risk, will live.

The pull request was perfect.

Green tests. Clean diff. A migration, a rollback path, the awkward edge cases I'd normally hand to a mid-level, all handled. I'd asked the agent for it an hour earlier, and there it was, better formatted than anything I write at 4pm on a Friday.

I went to approve it.

And my hand stopped on the button.

Because in the systems I work on, approving doesn't mean agreeing. It means owning. My name on the merge is me telling the business this is safe to ship. And staring at a flawless diff I had not written, I hit the uncomfortable part: the machine had done in an hour what used to take a week, and I still needed roughly that same week to actually trust it.

That distance has a name now. I read a paper that put a number on it. And once you see it, you can't unsee where the whole AI economy is heading.

Execution is free, verification isn't

Two Curves Headed in Opposite Directions

The cost of making things is falling off a cliff. Compute gets cheaper, models get sharper, and an agent will write the code, run the tests, draft the analysis, file the report, all at almost no marginal cost. Call that curve cA. It's diving.

The cost of checking those things is barely moving. Verification still runs on human time. Experience. Context. Slow feedback from the real world. Someone has to read it, understand it, and put their name on it. Call that curve cH. It's flat.

Draw both on the same chart and the story writes itself. One line collapses. The other holds. The space that opens between them is the whole game.

The paper calls it the Measurability Gap: the distance between what a machine can execute and what a human can actually verify. And here's the part that matters: it widens by default. Not because anything goes wrong. Because the two curves were never going to move at the same speed.

The Runaway Risk Zone

Sort any task by two questions. Is it cheap to automate? Is it cheap to verify? You get four boxes. Three of them are fine.

The fourth is the one that should keep you up at night.

Cheap to make, expensive to verify. The paper calls it the Runaway Risk Zone, and it's exactly where modern AI thrives. The agent produces something fast, polished, plausible, and checking whether it's actually correct, safe, and aligned costs more than the thing saved you. So nobody checks. The output looks like productivity. It registers as productivity. It just isn't, and you won't find out for a while.

That zone isn't a corner case. As compute grows, more tasks slide into "cheap to make." As deep experience thins out, fewer stay "cheap to verify." The zone is eating the map.

Forget "Routine vs Creative." The Axis Is Measurability.

We've been comforting ourselves with the wrong split. Machines take the routine, humans keep the creative. That was the story.

The real axis is measurable vs unmeasurable. Anything you can compress into a metric, a benchmark, a test, a feedback loop, the machine takes, however creative, however senior it looks. Writing code is measurable. Passing a bar exam is measurable. Diagnosing from a clean dataset is measurable.

Your credentials don't protect you. Your ability to operate where quality can't be reduced to a number does. That's the uncomfortable career memo hiding inside the paper I was reading: the premium isn't moving toward the qualified. It's moving toward the unmeasurable: judgment, intent, the calls nobody can score yet.

The Loop That Eats Its Own Tail

Here's why that famous "keep a human in the loop" line is a comforting lie. That loop is being corroded from two directions at once.

  • From below. Automate the junior work and you stop minting juniors. But junior work, the grind of doing it wrong and fixing it, is exactly how you grow the seniors who can verify the hard stuff later. Kill the on-ramp and you starve your own future supply of verifiers.
  • From inside. Every time an expert corrects the machine, labels the tricky case, explains the weird edge, they hand over their tacit knowledge as training data. The expert is feeding the system that shrinks their own scarcity. The paper calls it the Codifier's Curse, and it's a beautiful, brutal name for what happens.

So human-in-the-loop isn't an equilibrium. It's a phase. What you're left with is a sandwich: humans set intent at the top, machines fill the whole middle, and a thin, shrinking, human layer at the bottom verifies and absorbs the risk. Fewer people, holding more liability, watching more output go by.

The Trojan Horse

The danger was never obviously-wrong output. You catch that. The danger is plausible output: right on every metric you have, quietly violating the intentions you never thought to measure.

It ships clean. It books as activity. And it accrues hidden debt: the bug, the fragility, the biased call, the security hole, the decision that was optimized instead of correct. The invoice arrives later, when the feedback finally catches up, and by then it's expensive.

Scale that across an economy and you get what the paper calls the Hollow Economy: more measured activity, less real value, less human agency. Everything looks busy. Nothing is trustworthy.

The alternative, the Augmented Economy, isn't automatic. It only happens if we scale verification as aggressively as we've learned to scale generation. Observability. Provenance. Ground truth you actually own. The boring infrastructure of trust, treated as production capability instead of a compliance afterthought.

So, Now What?

If you're an engineer, the move is up the chain, and it's not optional. Out of pure execution. Toward intent, arbitration of trade-offs, verification, and a reputation you can actually stake your name on. It's a little uncomfortable to say, I know, but the truth is that execution is becoming a commodity.

If you run a company, the reframe is sharper: verification stops being a checkbox and becomes your moat. The advantage isn't generating more output with AI. Everyone will have that. It's generating output someone can trust, audit, and underwrite. That's a product.

A few months ago I wrote that your veto over what the machine sees is the job. This is the bigger version of the same sentence. Your veto over what the machine ships is about to be the only part of the job that pays.

Anyone can make the machine produce. The scarce, expensive, irreplaceable thing is the one human willing to look at what it made and say, with their name on it, whether they truly approve of it or not.