AI-Assisted Dev3 min read

All thirteen tasks passed review. The full review found three blockers.

I ran a code review at the end of every task. All of them passed. A single review over the whole thing found blockers — the defects lived between tasks, not inside them.

#methodology#verification#reality-check#gotchas
Concept diagram: per-task reviews each cover one output while the boundaries between them go unchecked
A diagram summarising the post.

I built an app as thirteen tasks and ran a code review at the end of each one. All thirteen passed.

Just before release I ran a full review over everything, out of habit. Three blockers.

What is the unit your reviews operate on?

Where the defects lived

All three were correct within their own task.

One task wrote the store description. It read well and was translated into 15 languages. Pass.

Another task filled in the data safety declaration. It was accurate. Pass.

That the two documents say opposite things was in neither task's scope.

The other two have the same shape: the task that described a feature was not the task that implemented it, and the task that counted permissions was not the task that added the billing SDK.

Why this is not bad luck

Splitting work splits the review context with it. A reviewer sees that task's diff and goal. Whether a later task broke an assumption an earlier task established is outside that field of view.

And the finer you slice, the more boundaries you create. Thirteen tasks means twelve boundaries.

A fork in the road

Two responses: make tasks bigger to reduce boundaries, or add a review that looks at boundaries.

Which one?

I took the second. Bigger tasks make reviews shallower and rollbacks harder. Instead I pinned a full review before release into the process.

It costs one review. This time it caught three blockers.

My own audit had the same hole

Before that full review I ran my own go/no-go audit: unit tests, lint, bundle verification. I called it "code GO."

During that audit I dirtied the device database and left the instrumentation suite red. The audit never re-ran instrumentation. It issued GO on unit tests alone.

So I wrote down one more rule:

Unit tests and lint do not entitle you to say "code GO." Include instrumentation, and put copy and declarations in one table.

Self-check

  • If you split work into N pieces, do you have a check that looks at the N−1 boundaries?
  • When a later stage changes an assumption an earlier stage made, is there a stage in your process that catches it?
  • Do your release gates cover artefacts outside the code — copy, forms, config, store assets?

The honest part

Running that full review was a habit, not a design. Without it all three would have shipped, and one of them was a sentence telling users something false.

This is not the first time either. Every app passed, then every app failed is the same family. The sum of individual checks is not a check of the whole.

Count how many units you split your last feature into, then count how many checks looked between them. Mine was thirteen to one.

Related