Skip to content
wildcard

Insights3 min read

Specs are the new code review

When agents write the code, humans add value in the spec and the evals. How Even's process runs: scope, build in parallel, review agent, evals, human gate.

Cliff des Ligneris

Specs are the only code review that scales to agents. The diff is too late.

When one engineer writes code, reading the pull request is a reasonable place to catch mistakes. The volume is human. When a team of agents writes it, the volume is not. On Even, a day of agent work produces more changes than I can read with care. If my value is in reading diffs, I am the bottleneck and the agents are waiting for me.

So the value moved. It moved up, into the spec, and down, into the evals. Here is how the process runs.

Scope. A goal becomes a set of small tasks, each with explicit acceptance criteria. I write the goal. Agents draft the breakdown. I approve it or send it back. This is where most of my judgement goes now. A vague task comes back as confident, wrong code. A precise one, with the invariant written out, comes back right more often than not. The spec for a slice says what must be true afterwards, which tables may change, which must not, and what the negative test looks like.

Build in parallel. Implementation agents work on separate branches at the same time, each with the codebase conventions and its task in context. I do not watch this part.

Review agent. Before a human sees anything, a second agent checks every change against the spec and the conventions. Only what passes reaches me. This is the step that caught two things on Even recently. The reviewer found that a cross-tenant test could pass on a leak because it committed after the check instead of before. It found a concurrent-write race on roster rows that no human had looked for. Neither was in my head. Both were in the spec as “must not”, and the reviewer took the spec seriously.

Evals. The operator’s behaviour is tested against a growing set of concrete situations with expected outcomes. A change that moves an eval the wrong way does not ship, however clean the diff is.

Human gate. I approve every release and every rule the operator may apply on its own. By this point I am reading a summary, the eval results and the reviewer’s notes, not every line.

What this means for a CTO with 30 to 300 engineers: your senior people are about to become spec writers and eval authors, whether you plan for it or not. The teams that struggle are the ones that keep review as the control point and watch the queue double. The teams that work move the control point earlier and make it testable.

Two things I would do this quarter. Write the acceptance criteria for one feature as a spec an agent could build from, and see what is missing. Count your evals, meaning behaviours a machine can check, and set a target.

Specs and evals are where the humans are. Where does this break? Show me a team that keeps up with agents by reading diffs.

Back to insights

Talk to us.

Thirty minutes. We talk about your situation and you leave with a next step or an honest no.