How I work with AI

Most of what gets built on my desk now is AI-assisted. Some of it is code I can read but couldn’t write. The way I see it, that only works one way — the discipline around the tool has to be stronger than the tool. Here’s what that discipline is, and the missteps that taught me each piece of it.

Frame the problem before picking the tool

Most bad AI output is a good answer to the wrong question.

A while back I had a tool watching job postings for me. Three weeks in, twenty-nine roles found, zero applications sent. So I sat down and wrote a page taking the tool apart: its scoring, its duplicates, the fact that it never pushed me to act. Sharp critique — wrong target. The tool had done its job. The reason nothing went out was upstream, and it was mine: the portfolio wasn’t ready and the résumé wasn’t focused. One question, why zero, would have found that in a minute. I asked it last.

What I took from it: a result at the end of a chain doesn’t tell you which link failed. Now the first question is always about the whole path, and the tool gets blamed last.

Sort claims by what it costs if they’re wrong

Most claims are cheap to be wrong about. Four aren’t:

  1. “Tests pass.”
  2. “This is the only place it occurs.”
  3. “This is the root cause.”
  4. “Verified in the browser.”

Those four carry a command and its output, or they’re labeled unproven and travel that way. Everything else can be wrong cheaply, and usually is, and that’s fine.

They’re on the list because they’re the four I’ve watched fail. Not one of them failed because the reasoning was bad.

A claim with no stated control isn’t evidence

I spent two days watching four models work real defects and graded every result. Going in, I expected the mistakes to be bad reasoning about code. Almost none were. The code reasoning was good.

What broke was the instrument. A file count that swept in hidden internals and came back a plausible number. A search through stylesheets for a style the framework had inlined somewhere else, reported as “the fix isn’t there.” Each one arrived as a finding with a number attached, and each one dies to a single question: what did you check that against?

What I took from it: a check has to be run against a case known present and a case known absent before its answer means anything. That’s the control. No control, no evidence, however confident the number looks.

Install the discipline instead of trusting the output

I can’t audit a diff line by line. What I can do is build the environment so the careless version of the work doesn’t pass.

What I mean by discipline is this: gates and guards that fire whether or not anyone remembers to check. A stop-gate that blocks a model’s first “done” and turns a read-through into real browser-and-test execution. A pre-execution guard that catches a masked exit code and teaches the shell-correct way to read one. Those aren’t better prompts. They’re the opposite of prompting, and they transfer to any model I run, including ones that have never seen the codebase.

Say what isn’t proven, first

Naming the gap yourself turns a reviewer’s catch into your disclosure. I’ve watched it read as strength every time, and I’ve never once seen the alternative work.

The missteps, and what each one changed

I keep a log of the times the work went wrong. It’s not there to be hard on myself. It’s there because the pattern only became visible at four or five entries, and I’d have sworn before that it was a thinking problem. It wasn’t. Six of the first seven entries were the check failing and reporting success.

The find-and-replace that did nothing. Setting margins in a generated document. The edit ran, returned clean, and the margins never moved, because the pattern had matched nothing and a zero-match substitution looks exactly like a successful one. Now: any edit to a file I didn’t author reads the artifact back afterward. Zero matches is a failure, not a no-op.

The search that couldn’t match. Checking a document for a string, twice, and getting “missing” both times. The file had escaped the character I was searching for, and the renderer had changed the case. Both searches were structurally unable to find what was there. Now: before reporting a string check against anything rendered or generated, say what a failing result would look like and confirm the check could produce it.

The page that hadn’t loaded. A capture read “absent” because the page was still hydrating. Now: query the element directly, or scroll first. One capture of a lazy page is not a read.

The scrub that passed. Cleaning a portfolio asset of names that weren’t mine to publish. Verified against 1,692 exact strings across the whole codebase. Zero hits. Clean. Then I opened the rendered page, and a name was sitting right there in a spelling variant none of the 1,692 covered. Now: the rendered surface is the gate. Grep is a pre-filter.

The deletion that never happened. Two true facts side by side: records that began in May, and a retention setting that deletes anything older than ninety days. I welded them into a cause, concluded months of history were gone, and said so urgently enough that someone acted on it inside a minute. One command falsified it: files well past the cutoff were still there. Someone had simply started using the tool more. Now: before reporting a number, say what it would look like if the instrument were wrong. And anything that arrives with “act on this today” gets one more check, precisely because it’s about to be acted on.

None of those raised an error. Every one handed back a clean, confident, wrong answer. That’s the shape of the risk with this kind of work, and it’s the same shape whether the decision was mine or the model’s. The fix was the same too: a check that cannot fail is not a check.

How these pieces get written

Since the question comes up: yes, these are AI-assisted, and they’re mine.

The ideas start as notes and as me talking them out loud. Then they go to the model, which drafts against my voice guide and my own recorded phrasings, and I read it back and say what’s wrong. That’s usually a lot. We go through several revisions until it says what I think, in a way I’d say it, with nothing in it I can’t stand behind. The missteps above are in here because they’re the part of the process I’d least like to publish, and that’s exactly why they belong.

I don’t want the robots writing for me. I want the robots repeating me.

Where this is going

I’m not the only one who’s going to work this way, and the discipline is the transferable part. The point of writing it down is so the people I work with can pick it up: the developer who’s going to review the code, the QA lead who needs to test it without reading it, the product owner who needs to know what’s proven and what isn’t.

That’s the job I want. Get on the bus. We’re going for a ride.