Automation Pipeline4 min read

My Scheduler Was Only Running the Publishing Half

A publishing job ran every evening on a schedule. The queue it pulled from, though, still had to be filled by hand. This is the record of wrapping generation and publishing in one script and pointing the scheduler at it.

#launchd#automation#scheduler#shell#content-rotation
Left: the scheduler calls only the publish script while the queue is filled by hand. Right: the scheduler runs one wrapper script that calls generation, then publishing.
Automated publishing didn't mean an automated pipeline. The scheduler had to call a different entry point.

There was a job that ran every day

My short-form video pipeline had a publishing script that the scheduler called every evening. macOS launchd ran daily_publish at a fixed time, and daily_publish took items off the publish queue and posted them. That part was automated.

Filling the queue was a separate job. rotation.py captures web app result pages in order and adds them to the queue. I had already built that engine and had it save its rotation position to a file on disk. But when I looked at the scheduler config again, launchd called only daily_publish. The generation step was not on any schedule.

In your own automation, what entry point does the scheduler actually call, and who produces the input that entry point consumes?

"Publishing is automated" and "the pipeline is automated" are different claims. If only publishing is scheduled, filling the queue falls to a person, or to whatever agent session happens to be open. Once that person stops, the queue only gets shorter.

Why it was easy to miss

The backlog already in the queue acted as a cushion. With generation stopped, publishing still ran normally for a while, and from the scheduler's side every day looked like a success. If you only watch whether the publish job succeeds, you never see that the step producing its input isn't scheduled. The problem would only show up on the day the cushion ran out.

What would you do: put generation inside the publish job, or add a new outer script that wraps both?

One wrapper, and a different scheduler target

I went with a thin outer script. engine/autopilot.sh is 22 lines and does two things: it runs rotation.py to generate the day's queue items, then runs daily_publish to post. I didn't change either of the existing scripts.

Then I changed the launchd plist so it runs autopilot instead of daily_publish. That was a 2-line change. It runs at 18:30. The same commit also added one line to .gitignore.

Because the rotation creates a new theme every day, this setup doesn't need anyone to open a session or step in to keep the queue filled. That's what "no intervention or session needed" means in the commit message.

Verified without actually publishing

I didn't want the first test of an unattended script to be a real publish, so I added a DRY=1 mode to autopilot.sh. I checked three things:

  • That the chain runs end to end, from generation to publishing, in DRY mode
  • That Chrome capture works in a clean environment
  • That the queue still has a backlog cushion

The second one matters most under launchd. A process launched by launchd doesn't get the same environment as my terminal shell. Capture that works fine in a terminal can fail under the scheduler, so I separately checked that it works with the shell environment stripped.

Self-check

  • Have you opened the scheduler config and confirmed that the entry point it actually runs covers the pipeline from start to finish?
  • If generation stops while publishing keeps succeeding, would anything tell you?
  • Have you run your unattended entry point end to end with no shell environment, as a dry run with no real side effects?

The honest part

As of this commit, I had verified the DRY chain, clean-env capture, and that a backlog was still in the queue. This record doesn't show launchd actually calling autopilot at 18:30 and publishing on its own. The commit title says generate → publish → analyze, but the body describes only generation and publishing, so I can't say here how the analysis step is wired in. This record also doesn't tell me what the publish step does on a day when generation fails. I'm hoping the backlog covers those days, but I haven't checked.

Related