Dan North has a pattern called spike and stabilize. Instead of deciding up front whether you’re writing a throwaway prototype or production code, which means deciding at the moment you know the least, you spike to learn, then stabilize what turned out to be worth keeping. Liz Keogh has written a lot of the clearest material on it, and North’s own talk on patterns of effective delivery is where it starts.
I’ve been using it for years. It’s the honest version of how good work actually happens.
It also hand-waves in two specific places, and I’ve spent the last year building on top of both.
What a spike gives you before any of that
Two things, every time, and neither is in the pattern as written.
You learn the codebase by building inside it. A spike that lives in the real repository has to find where the new thing plugs in, what it can reuse, and what it would break. There’s no faster way to understand an existing system than trying to put something new into it. By the end I know the seams, and I know them because I hit them.
You get to use the thing. A spike completes the experience, so you can walk the journey in the real product or a stand-in for it, with realistic data, at speed. That’s a different kind of knowing from a mockup. Half of what I change after a spike, I’d never have seen in a static screen.
Those two are why I spike at all. The rest of this is about making what comes out of it trustworthy.
Soft spot one: “measure, don’t trust feeling”
The pattern says to measure whether the spike is worth stabilizing. It doesn’t say how.
In practice, “measure” collapses into someone senior looking at the prototype next to the production build and saying it looks right. I’ve been that someone. I’m reasonably good at it. I’ve also been wrong, and the wrongness doesn’t show up until QA, or a brand switch, or a viewport nobody checked.
So I made it mechanical. A parity gate diffs the prototype against production component by component, per interaction state. Structural components node by node, canvas-rendered ones visually at the wrapper. A change either matches the frozen answer key or it fails, loudly.
That’s the whole move: turn a principle into a gate. “It looks right” becomes “it matched, or here’s the delta.”
Soft spot two: the spike is one-shot
Classic spike-and-stabilize freezes once and harvests once. It’s quiet about what happens at the next release, which is where prototype-to-production drift actually comes from.
So the spike stopped being disposable. I restructure it to mirror the production component tree one to one, one prototype fragment to one production component, indexed by a manifest, and then keep it alive. It co-evolves, release over release. A twin — not a sketch.
The manifest matters more than it sounds like it should. The mapping between prototype and production has to live somewhere other than a person’s head, because a head is exactly where drift comes from.
The loop
Carve from the current build. Dial in the change on the twin. Get sign-off. Gate it. Roll up to production. Repeat.
Two rules make it work:
Carve, don’t rewrite. Cut the prototype along production’s existing component seams. Adopt them. Don’t invent better ones, however tempting. The moment your seams differ from theirs, the twin stops being a twin.
Gate both ends. The carve has to come back zero-delta against the frozen baseline. The roll-up has to match the production render. One gate catches a bad cut; the other catches a bad landing.
It’s stack-agnostic. Any component framework, any framework-free prototype, any automated browser diffing.
Where the AI actually sits
This is where I’ll be specific, because vague AI claims deserve the skepticism they get.
AI does the front half: the mechanical carve, the port, running the checks, generating synthetic data at realistic volume. That’s genuinely a lot of the labor, and it’s the part that used to make this approach too expensive to bother with.
The judgment stays mine. What’s worth stabilizing. What the delta means. Whether a failure is the gate being wrong or the work being wrong, which is a real question and the gate can’t answer it.
That division is why the method works now and didn’t five years ago. Not because the tooling got smart — because the tedious half got cheap.
Two questions this raises, answered elsewhere
A spike that lives in the real repo produces a real diff, and nobody on a dev team wants to read a thousand files of changes. How you hand that over without burying people is its own problem, and I’ve written it up separately: cut the backlog from the same source as the build.
And the obvious follow-up, how big should a spike be, gets its own piece too, because the answer is “smaller than it looks finished,” and that needs more room than a paragraph: how big should a spike be.
What I’ll tell you that the pattern won’t
It’s being proved out. I’m not selling a framework.
The production-side check is human-supervised, with auth, tooling, and a real browser. It isn’t push-button. I’d love to claim otherwise and I’m not going to.
And the value isn’t that it guarantees parity. It’s that parity becomes provable, so the conversation stops being about whose eye is better. That’s a smaller promise than most process advice makes. It’s also one I can show you the numbers on.
