Automation Pipeline5 min read

My sync job reported zero gaps, and it was telling the truth

A job wires store links into my pages whenever an app passes review. It reported zero gaps three runs in a row. Three landing pages still pointed at an App Store search. The job wasn't wrong. Its list was short.

#automation#app-store-connect#gotchas#data-quality#reality-check
Left: a sync job that diffed six surfaces and reported zero gaps. Right: a seventh surface it never listed, a landing page still carrying an App Store search link
A set difference is zero only inside the list. It never asks about what's outside.

Passing review fires no event

I run 59 iOS apps and 37 Android apps on my own. When one passes review, a lot of places need a store link: the app directory, the home page in four languages, the sitemap, the lists my promo bots read, platform badges, the store monitor.

The catch is that passing review sends me no signal at all. One app actually sat on my home page and promo lists as "Android only" for more than a week after its iOS version went live.

So I built a job that runs every six hours. It is simple:

  • Measure the set of public apps on the stores (App Store Connect plus the public lookup API, and Play listings).
  • For each surface, read the set of apps it links to, straight from the live page.
  • Compute public set − surface set, per surface.
  • If anything is missing, an agent wires it in, and then the job runs the diff again to decide. The agent saying "done" counts for nothing.

Yesterday the job printed gaps 0 three runs in a row.

I did one more sweep by hand, with a different question. Not "are there gaps where the job looks?" but "where on disk is an app ID written down at all?"

That turned up a surface the job had never read: per-app landing pages. Across three apps, nine files (the landing page plus privacy, terms and support) had an App Store button like this:

https://apps.apple.com/search?term=Somnul

It was a placeholder from building the page before launch, when there was no app ID yet. It survived the launch. Tapping it lands you on search results, not the app, and competing apps show up right next to it.

The job wasn't wrong. It was exact about the six surfaces it knew. Landing pages just weren't on its list.

When your sync or audit tool says "0", which list was it counting over?

The same sweep found one more

While scanning, I reopened the generator for my promo video app list. It had 44 apps hardcoded. The file it supposedly generates had 58, because recent apps had been added to the output directly.

Run that generator once and 14 apps, with their hand-written copy, silently disappear from the output. You might notice the count drop, but not if another app was added the same day.

Would you refill the generator with the 14, or change the generator?

I didn't refill it. The hand copy and icon paths had already drifted from the generator's rules, so a rerun would still overwrite them. Instead, it now counts the apps a rewrite would lose and aborts if there are any. Run it today and it prints the 14 names and exits with rc=1. The output file doesn't change by a byte.

A seventh surface, used the very next day

I added landing pages as the job's seventh surface. It checks that each public iOS app ID and Play package appears on the landing page as a direct link. The regression test now includes a search?term= landing page that must count as a gap.

This afternoon an app (Mosslock) passed review. The 15:16 run found 6 gaps: the directory, the home page in four languages, and the promo list. The agent wired them in 13 minutes, and the re-measured diff was 0.

But the landing page wasn't among those 6, even though it had no store link at all. The job finds landing URLs through the directory's links, and the new app wasn't in the directory yet, so the first pass never read its landing page. The agent added the button anyway because its runbook says to. The re-measure after the directory was wired did read the landing page, and that's when 0 came out. One day later I was looking at the same lesson again: even the seventh surface depends on where the list comes from. So landing URLs are now also derived from the store's bundle IDs. A new app may not be in any of my lists yet, but it is on the store.

Build the list with grep, not memory

What helped this time wasn't a new check. It was how the list gets built. Instead of writing down the surfaces I remember, I searched the whole disk for files containing ten or more app IDs. Lining that result up against the surface list exposes catalogs the list doesn't have.

The same scan works in reverse: dead links to apps that have been pulled. That came back at 0.

Self-check

  • Where did the target list of your sync or audit job come from? If a person wrote it, have you ever compared it against a list built by grepping for identifiers?
  • Is there a step that goes back for the placeholders you set before launch (search links, stand-ins, "coming soon") after the launch?
  • If you hand-edited a derived file, do you know what would disappear if you ran its generator right now?

There's a related story: two weeks of green jobs and a stale site. That time a success signal stood in for the result. This time the number 0 stood in for the scope.

The honest part

The seventh surface was found by a person sweeping by hand, not by the job. I have no evidence there isn't an eighth. The repeatable part is the grep scan, and I haven't scheduled it yet. I now read "gaps 0" as "0 across seven places." It's worth counting how many places your own 0 covers.

Related