One more daily job
I have a pipeline that turns my dev blog posts into short videos. Until now I wrote the scripts by hand. Rendering and the publish queue were already automated. This change automates the front of the pipeline too.
The flow is simple. Pick a blog post that hasn't been converted yet, hand it to an LLM CLI to turn into a script JSON, and if it passes a structure check, render it and rebuild the publish queue. It runs daily at 10:00 via macOS launchd. Publishing happens at 11:00, so generation is scheduled an hour ahead.
Up to this point it's an ordinary automation. The question I spent the most time on was different: when should this job stop?
Try it on your own project. If you run an automation that produces something every day, is its input finite, or is it the kind an LLM can generate forever?
The easy path was infinite generation
If an LLM writes the scripts, you don't strictly need source material. Feed it a list of topics and scripts keep coming. The queue never runs dry.
The problem is what those scripts rest on. The videos on this channel summarize blog posts about things I actually went through. A script generated without a source post has nothing behind it. The format would look the same, but events I never went through and numbers I never measured could slip in.
What would you do? Generate more so the queue never empties, or accept that some days it will be empty?
I tied the source to real posts
I chose the second option. The only source for the generator is blog posts that were actually published. At the time of the commit, that was 27 posts. The generator picks one that hasn't been converted, and if none are left, it produces nothing. No separate stop condition needed; it stops by construction when the source runs out.
I also set the order. Each post has a hook type, and the generator converts higher-priority hook types first. If the supply is finite, it's worth deciding what gets used first.
The JSON the LLM produces isn't rendered directly. It goes through a structure check first, and malformed scripts never reach rendering. Only scripts that pass are rendered, and then the Korean and English publish queues are rebuilt.
The part with no cheap fix
This design has a cost. With 27 source posts, the videos end somewhere around 27. After that, the only way to fill the queue is to write more posts. The automation didn't remove the bottleneck. It moved the bottleneck to the speed at which I write.
No setting or switch fixes that. Switching to infinite generation would keep the queue full, but it would break the premise that this channel is built on real records. I chose to let the queue run dry.
Self-check
- Can you say in one sentence whether your daily generation job's input is finite, and what the job does when it runs out?
- Is there a gate that checks the format of LLM output before it moves to the next step (rendering, publishing)?
- Is there enough slack between your generation job and your publish job for generation to finish?
The honest part
This post is based on a single commit that added the automation. I don't yet know how the generated scripts compare to hand-written ones, or how these videos performed. The structure check validates the JSON's shape; it doesn't guarantee the content matches the source post. I haven't verified that the one hour between 10:00 and 11:00 is always enough for generation and rendering. And I haven't yet seen what the pipeline looks like on the day the queue is empty after all 27 posts are used.