Skip to content
All posts

6 min readAI, Delivery

Tests first, checks not prose: how we build with AI now

AI is not magic, and probability is a risky way to write code. Five rules for building with AI agents, all enforced by machines, and what each one prevents.

By Dave Tormey

A laptop with code open, standing in for tests written before the implementation

AI is not magic. A language model writes code by predicting what is most likely to come next, and using probability to write code invites risk. The answer is not to stop using it. It is to put the decisions somewhere the model cannot reach. This is the method I use now, for my own product and for clients, and it replaces the methodology I published in April.

The principle

The model's job is typing, not deciding. Everything that decides, what gets built, whether it is right, whether it merges, sits outside the model, in a plan a person approved or a check the model cannot alter.

The five rules

1. Plan before code. Every module has a written plan before any code exists: its purpose in one sentence, its public interface, what it trusts and what it does not, every case where it must refuse and the exact error it returns, and a test list of one line per test. The owner approves the plan. If the plan cannot be written as a test list, it is not finished.

2. Tests before implementation. The tests from the plan are written first, committed on their own, and seen to fail for the right reason. Then the implementation makes them pass without changing them. If a test looks wrong, the agent stops and says so. It may not edit a test to make it pass. The reviewer later checks that no test changed after it was committed as failing.

3. A rule that is not checked does not exist. Every rule lives in one command that runs on every commit: file size limits, no pattern matching in production code, no module reaching into another's internals, only one module allowed to read the environment, no swallowed exceptions. If a rule cannot be expressed as a check, it is deleted rather than written down. The agents read one short instruction file, and it says almost nothing the checks do not enforce.

4. Each agent works on one thing, on its own short-lived branch. Agents work in parallel only on pieces whose boundaries are already fixed, so they cannot trip over each other.

5. Nothing merges without independent review and a human's button. An agent that did not write the code reviews every change against the plan and the tests and answers five plain questions. The owner reads a one-page summary and presses merge. A change too big to review in ten minutes is too big.

What each one prevents

Plan-first prevents the agent deciding scope. Tests-first prevents the agent grading its own homework. Checks-not-prose prevents drift, because a check cannot be talked round and a paragraph can. Independent review prevents a quiet failure common to AI-built code: code that passes every test because the same agent wrote both.

The harness pattern, kept and fixed

Anthropic has published a pattern for long-running agents: a feature list, a progress log, a start-up script, one piece of work per session. It is a good pattern, with one weak point: the agent can mark its own work as passing. Here the feature list is the test list, turned into real failing tests before any code; "passing" is computed by the checks; the progress log is the commit history. Same shape, with the honesty moved out of the agent's hands.

The numbers

Three, every Friday, from a script: modules built over planned, tests passing over total, real workloads passing end to end over total. Nothing merges that makes any of them worse. A daily independent audit reads the commits and reports what it finds as an issue. It never edits anything.

Does it slow things down?

Yes, at the start. Writing the plan and the failing tests before any implementation takes longer than typing straight at the problem.

The honest comparison is not speed. It is whether, at the end of the day, the owner can read every line that landed and would sign their name to it. This method makes that possible, and it is the only measure that survives contact with a customer.

Dave Tormey

Dave runs TAGD Ventures in Brisbane. He has spent more than twenty years building and running software for government and regulated industry, including more than a decade as CTO, CIO and CISO of a software company serving law enforcement.

Get in touch