Shipping & Infra6 min read

I excluded my own expansion to protect the control group — and then nothing was judging it

To keep one verdict clean, I removed newly expanded apps from the control group. That was correct. What I missed is that those apps then appeared in no verdict at all — exclusion prevents contamination and opens a monitoring hole in the same move. Every exclusion needed a row of its own.

#gates#preregistration#monitoring#measurement#reality-check
Left: a newly expanded app left inside the control group inflates the expected value, so the treatment group's bar rises on its own. Right: removing it clears the contamination but leaves that expansion inside no verdict at all, a blind spot
Exclusion is half the job. The other half is that expansion's own row.

Last month I shipped 227 new pages and preregistered a gate for them. Verdict date 10-20, and the comparison is "the trend among apps I did nothing to" — because the whole channel is growing, and absolute totals therefore say nothing (my control group grew 4.5x without me).

Since then I've added pages to two more apps. Per the rule, I removed those two from the control group. That part I got right.

Three days later I noticed the rest of it. Once removed, those expansions were being judged by nothing.

A question for you. The data you excluded from your experiment as contamination — where is it being judged now?

Why the removal is correct

The second metric looks like this:

expected = 32 × (control_group_in_window / 141)
PASS  ⇔  treatment ≥ 1.30 × expected

The control group is the denominator. Leave a newly expanded app inside it and the control grows, expected grows, and the bar the treatment group has to clear rises by itself.

That is the same class of contamination as moving the bar after seeing data, just pointed the other way. Do nothing and my own test gets harder as time passes. Structurally identical to waiting two days and finding my bar had hardened on its own.

So the exclusion list lives in code:

EXPANDED_HOSTS = {...}   # apps that received their own expansion inside the window come out of both sides

The baseline decomposition moves with it. If a newly excluded app was already earning long-tail traffic during the baseline window, its share moves from control to expanded. The total doesn't change — still 173. A self-check asserts that invariant. For the other app there was nothing to move: its measured baseline long-tail was 0.

That's the contamination guard. It works.

What the removal does next

Being out of the 227-page verdict means that expansion's effect will not appear in the 10-20 result. Of course. That is the entire point of excluding it.

The problem is what comes after. There was no other verdict containing it.

  • the 227-page verdict: excluded
  • its own verdict: none

So across shipping 25 pages, growing that to 26, and then shipping 13 more, not one of them had a written answer to "when is this a success, and when do I stop?" A job reads that roster every day; with no row to read, it stays quiet forever. Not green — not an item at all.

This is a trap I keep stepping into. When my monitor was watching half the fleet, and when it was guessing the names of what to watch, the symptom was the same. Missing things don't show up red. They show up as no color.

What would you do?

Two ways out.

  1. Put the excluded apps back into the 227-page verdict, so at least something judges them.
  2. Keep the exclusion, and give every exclusion its own row.

Option 1 is not available. It restores the contamination, and it mixes two expansions' effects into the 227-page number, so even a PASS tells me nothing about cause. A verdict you can't learn from when it passes isn't a verdict.

So, option 2, and one new rule: if you take something out of the control group, you create its own verdict row in the same commit. Exclusion and own-row are one unit of work.

The own row is two lines

Each expansion gets two entries.

futility   10-28   stop if indexed URLs == 0
verdict    12-26   long-tail sessions >= 6  AND  indexed >= 5/26

The important part: "not indexed" is never written as FAIL. It returns "cannot judge." If the pages were never indexed and I stamp FAIL from the session count, I haven't learned "this expansion approach doesn't work" — I've recorded 'these pages never existed in search results' as an expansion failure. Same mistake as a gate with no denominator, rejecting the asset when what's missing is distribution.

That's why the futility row's conclusion text is pre-written as "never indexed" rather than "expansion failed." Choose the wording on verdict day and the day's mood gets into it.

Put dates in the date column

One more trap rides along with this. The automation that reads the roster scans only the date column. But while explaining a rule, it's natural to write a sentence like "revisit this on 09-29" inside the body cell.

That date never fires. A human reading the table sees it; the machine does not. So there's now a check that blocks a commit when a future date appears in a rule cell — because one judgment date really was buried that way, and nobody would have looked that day.

Writing it in the roster and having the monitor see it are different things.

Three checks for your own work

  • Open the exclusion list of an experiment you're running right now. How many of those names have a separate verdict attached?
  • Does your gate distinguish "the denominator never opened" from "the result was bad" — or are both FAIL?
  • Is the document holding your verdict dates read by a machine? If so, which column does it read? Are there dates outside that column?

Conclusion

Contamination guards are almost always subtractive. And the moment you subtract something, it appears in no table. Not as a failure, not as a success. It's simply absent.

So exclusion is a two-step job now: remove it, and create the removed thing's own row. Do only the first step and you trade one clean verdict for one blind spot.

The honest part: the hole was open for three days, and what found it was not an automated check — it was me writing the next expansion's notes. The "excluded means own row" rule currently lives only in prose. There is a check that blocks a commit when a new row is left blank, but nothing yet catches adding a name to the exclusion list without adding its row. I'll wire that the next time I step on it. One incident, one check.

Do one thing today. Open your exclusion list and write one line per name saying where that name gets judged. The names you can't finish are your holes.

Related