Directing code I couldn’t write

I ship AI-written code into a repo that senior engineers review. I can read it — I couldn’t have written it.

That sentence makes people uncomfortable, so let me make it worse: I think that’s a normal condition now, and most advice about AI-assisted development assumes it away.

To be exact about where I sit: HTML and CSS, I’m a Jedi. TypeScript and the full stack, I’m a junior at best. I can follow the logic in the code that comes back and I can tell you what it should do. I can’t produce it on my own. So the work is describing what I want, pointing at the existing codebase as the guide, and dialing in until what comes out is accurate.

The good material on AI-assisted development, Simon Willison’s ongoing writing and the Fowler-adjacent work on verification speed, is written for a developer who can review the output. That’s the right audience and I learn from it constantly. But there’s a second audience nobody writes for: the person directing the work who cannot personally distinguish a correct claim from a confident wrong one, while the reviewers downstream absolutely can.

That’s me. Possibly it’s you.

The problem isn’t the code

When I actually measured it, the result surprised me.

I spent two days watching four models work real defects in a production codebase and graded every result. I expected the near-misses to be bad reasoning about code. Almost none of them were. The models are good at code.

What broke, every time, was the instrument, the check the model used to verify its own claim:

  • A file count that included .git internals. Came back as a plausible number.
  • A process check where the grep matched its own pattern in the process list.
  • A search through .css files for a style the framework inlines into JS chunks. Zero hits, reported as “the fix isn’t in the artifact.”
  • A conclusion about which session did what, drawn from a log written by a hook rather than the transcript.

Every single one arrived as a finding with a number attached. None arrived as an error.

And it isn’t a model problem I can wait out, because it happens to me on my own work. I was exploring a layout idea in code recently, a small thing, a column that needed to hold its width through a resize. I asked for the fix, got it, and asked for proof. Back came a tidy report: the offending style found, removed, build clean. I believed it. Then I dragged the window and watched the column ratchet wider on every pull, exactly as before. The search had looked in the stylesheet. The framework had put that style somewhere else entirely, and the report had never been capable of finding it.

The instrument passed. The thing was still broken — and only using it caught it.

The one tell that works

Every failure above dies instantly to a single question: what did you check that against?

Where a control was run, a case known present and a case known absent, the error got caught before it shipped. Every time. Where no control was stated, the number was a coin flip that looked like a fact.

So that’s the tell. Not confidence, not detail, not how good the explanation sounds. A claim with no stated control.

I can’t audit a diff. I can absolutely ask what the control was, and I can tell when there wasn’t one.

Four gates

I’m not the verifier. I route verification, to tools and to people who can actually do it, and I make sure nothing reaches a reviewer without evidence attached.

  1. Sort claims by blast radius. Most claims are harmless. Four are not, and they’re exactly the four that failed above: tests pass · this is the only place it occurs · this is the root cause · verified in the browser. Everything else can be wrong cheaply.
  2. Dangerous claims carry a command and its output, or they’re labeled unproven. A command and an exit code are readable without knowing the framework. “I checked” isn’t evidence. EXIT:0 sitting next to the command that produced it is.
  3. Demand the negative case. “It works” is half a result. The other half is the before-state, or the case where it fails. Pre-fix the column ratchets 342 → 630 → 724px; post-fix it pins at 436.4px, tracks a real resize, and returns. A non-programmer can check that shape. A broken instrument can’t fake it.
  4. Someone who isn’t the implementer checks first. Before the reviewer, always. The tooling for this was never the gap. Consistency of trigger was.

The part that’s actually leverage

I can’t verify the code. What I can do is install the discipline, and the discipline transfers to any model I run, including ones that have never seen the codebase.

What I mean by discipline is this: an environment where the careless version of the work doesn’t pass, built out of gates and guards that fire whether or not anyone remembers to check. A stop-gate blocked a model’s first “done” and turned a read-through into real browser-plus-test execution. A pre-execution guard caught a masked exit code and taught the shell-correct way to read one. A post-execution reminder produced consistent artifact-checking instead of trusting a green build message.

None of that is prompting better. It’s the opposite of prompting. And it’s available to someone who can’t write the diff, which is why I think it’s the highest-leverage thing in this whole practice.

What I don’t know

Whether the strong models are intrinsically careful or just compliant with that coaching. I’ve watched both explanations fit the data.

It matters, because those imply opposite investments. If they’re careful, the scaffolding is a cost I can shed. If they’re compliant, the scaffolding is the product and I should build more of it.

I don’t have that answer yet. I’m designing the run that would settle it.

Looking credible in front of people who can read the code

Credibility here is the relationship you have with the team, and I’m not afraid to come to the table. I’m not pretending to be a developer and never will. But I know enough that it isn’t over my head. I can participate, and I can shape the logic by thinking pragmatically about what the user experience needs to do.

Some habits, learned the hard way:

  1. Don’t go in expecting to be one of them. That’s most of what makes you credible. Bring your own strengths, name the ones you don’t have, and let the people who can read the code read it.
  2. Explain the logic and the use case, not the code. If I can walk a developer through the reasoning I used to get somewhere, and give them the situation it’s for, we’re speaking the same language even though I can’t write theirs.
  3. Hand over the spiked branch as the prototype. The experience and the journey are already resolved in it, so they can read the logic of what it should do straight from the working thing. How the code got there is whatever it took. Their job is to make it fit.
  4. If they don’t get it, the explanation is wrong, not the developer. When I can’t communicate what needs to get done, I go back and rethink how I’m explaining it and what the logic actually is.
  5. Share the evidence line, not the conclusion. “Fixed the sticky offset” invites interrogation. “Removed the two inline writes that clobber it, build exit 0, declaration confirmed in the built chunk, before/after computed values attached” ends the conversation.
  6. Say what isn’t proven, first. Naming the gap yourself converts a reviewer’s catch into your disclosure. I’ve watched it read as strength every time.
  7. Know what your reviewer actually enforces. Code that already matches a specific reviewer’s standards passes first-pass. This is researchable.
  8. Never carry a claim you can’t source. If the honest answer to “how do you know?” is “the model said so,” it isn’t ready to share.

That last one is the whole discipline, really. Being wrong in private is free. Being wrong in front of the people who can read the code is the cost function, and it’s a fair one.

Two things I’ve left out on purpose

Most of what’s above is me working with AI and the developers not. That asymmetry is the real story. If they worked with it too, or set up the context for me so a spike came back closer to what they’d want, I’d be handing over what they expect instead of what I could reach. I don’t know how to write that piece yet, but it’s coming.

And development is only the first channel. Without development, we don’t see anything. But QA needs to test the thing without reading the code, product needs to know what’s proven, and leadership needs to understand what it’s doing at all. It takes a village, and that’s a piece of its own.