One unattended job produced zero output for three days.
It never crashed, never alerted, and the dashboard stayed green the whole time. I found it this morning while investigating something completely unrelated.
What tells you about the day your automation ran and made nothing?
The log looked like this
09-15 18:05 social-growth-engine claude_fail
09-15 18:06 social-growth-engine claude_fail
09-16 10:00 social-growth-engine claude_fail
09-16 10:01 social-growth-engine claude_fail
09-16 18:05 social-growth-engine claude_fail
09-16 18:06 social-growth-engine claude_failThree runs, two retries each, all failing on the same item.
That job's queue held more than forty unprocessed items. None of them were touched for three days.
The culprit was one line
cands = pending(lang)
slug, md = cands[0] # only the firstIt takes only the head of the queue. If that fails, the run is over. The next run reads the queue again, the failed item is still at the front, so it picks the same one.
A single point of failure. No matter how much healthy inventory sits behind it, one dead item at the head means the whole pipeline outputs zero.
So why did nobody notice?
This is the real lesson. Two screens watch this job, and both of them looked fine.
First, the dashboard watches artifact freshness. Exit codes lie, so we decided to watch artifacts instead — a good design. But the artifact registered for this job is a log file. Every failure appends a line, so the log keeps growing. By freshness alone, this bot just did some work.
Second, the queue was full. Twelve items had arrived the same day through a different path (a manual batch). Queue depth looked healthy.
So "failure lines grow the log, therefore green" overlapped with "another path filled the queue, therefore green." Both instruments did their job honestly. Neither could see the truth.
What would you fix first?
The failing item, or the structure?
I ran that item by hand first. It succeeded in 22 seconds. Nothing was wrong with the item itself; what blocked it was a separate auth problem, covered in another post.
So fixing just that item would have restored today. And the next time any item gets temporarily stuck, I lose three days again. That is treating the symptom.
The structural fix:
for slug, md in cands[:MAX_CANDS]: # up to 3 from the front
got, generated = _gen_slug(lang, slug, md)
if generated:
return gotOne more trap right here
Implementing "on failure, move to the next" naively creates a new incident.
This pipeline has two stages: generation (an expensive call) and rendering (turning that output into a video). If a render failure also advances to the next candidate, every run burns three items of inventory. The script was produced fine; only the video failed. Yet the item gets marked done and the next item is consumed too.
So the return value splits in two — was something generated (decides whether to advance) and did it finish (the run's result). That boundary is what keeps a downstream failure from eating upstream inventory.
The tests aim at exactly that boundary. One asserts a run isn't empty when a candidate is blocked. The other asserts the call count never exceeds one when rendering fails.
Three things to check in your own setup
- Does your queue consumer take only the first item?
- Is that job's "proof of work" file one that grows even on failure? (Registering a log as the artifact usually means yes.)
- Do you have anything that reports a run that produced nothing while inventory existed? Freshness will never catch it.
The honest part
I haven't built #3. Today I fixed one single point of failure; there is still no watcher for "inventory available, output zero." Three candidates now have to fail in a row before a run comes back empty, which lowers the odds a lot — but lower odds and visibility are not the same thing.
And honestly, three days of loss was small here: 21 days of inventory downstream meant nothing starved. That is exactly what made it dangerous. The breakages that hurt nobody live the longest.
Pick one of your jobs and check what its last run actually produced — not whether the log grew.