Automation Pipeline3 min read

Half of My New Channel Spec Was About Not Touching the Old Pipeline

I designed new Shorts channels to drive app downloads. Before the feature design, I wrote the invariants that protect the pipeline already running. This covers the separate queue, token and channel, the flag branch, and the 48-hour verification gate.

#shorts#pipeline#isolation#feature-flag#verification-gate
Left-right contrast image showing new channels isolated from the existing pipeline by invariants
The longest part of the new channel spec was the isolation rules protecting the existing pipeline.

I opened the spec and 'what not to touch' was longer than 'what to build'

What I wrote this time is a design document, not code. It plans new Shorts channels that push viewers toward app downloads. The format is simple: a hook at the start of the video and a kick at the end. I open three new channels, in Korean, English and Japanese, and App Store native apps come first. The kick is a CTA that sends the viewer to the profile and then to an app list page, with a retention loop that brings them back.

While writing, the document tilted somewhere else. The longest part was not what the new channels would show. It was how to leave the existing pipeline, which already runs every day, unharmed.

The most common accident when attaching a new feature to a running system is not a bug in the feature. It is the feature touching something shared, so the part that worked fine stops with it.

So here is a question. If you attach a new branch next to automation that already works, what guarantees that the old side keeps running when the new side fails?

The six invariants I wrote

The spec lists six invariants that isolate the running pipeline. The core ones:

  • Separate the queue, so the new channels' jobs never mix into the existing queue.
  • Separate the token, so an auth problem on a new channel does not spread to uploads on the existing one.
  • Separate the channel: the three new ko/en/ja channels are independent of the existing ones.
  • Branch the render step with a render_hook flag. With the flag off, the existing path runs unchanged.

The remaining invariants are details that keep those four from being broken, so I won't expand on them here.

Finishing the infrastructure does not mean switching it on

The other piece is order. I finish the infrastructure first, and only after a 48-hour verification gate does autonomous drip, meaning automatic posting, begin. Right after building, all I have is a belief that it will probably work. I want to run it for real, confirm the old side is fine, and only then hand over to autonomous operation.

What would you check first before switching on a new branch? I decided to check 'is the old side unchanged' first. That is why isolation is fixed into the structure and the switch-on moment sits behind a gate.

Self-check list

  • Does the new feature share any queue, token or output target with the existing system?
  • Is there a flag that turns the new code path off, and have you confirmed the old path behaves identically with it off?
  • Before handing the new branch to autonomous operation, have you set a verification window that shows the old side is still fine?

The honest part

This post records a design document, not results. I don't yet know whether the 48-hour gate passed, how much traffic the new channels actually send to the app list page, or whether the retention loop works. I also haven't confirmed that the six invariants will stop every kind of accident in production. A spec only states intent, and the gate and the operating data will tell me whether that intent was right.

Related