The feature was two lines
The job was to port a comment-handling script I already ran on Threads to four Instagram accounts. What it does is simple. It sorts incoming comments into spam/abuse and normal. Spam and abuse get hidden, and normal comments get a reply written by Claude.
That takes two lines to describe, but the commit came to 216 lines of Python, 45 lines of tests, and a 36-line launchd plist. When I reread the 216 lines, only part of them did classifying or reply generation. Most of them were code that makes the bot not act when it shouldn't.
In your own automation, which part is longer: the code that does the job, or the code that stops it from running at the wrong time?
This bot writes in public and hides what other people wrote. If it misfires, the damage is hard to quietly undo. So this post is less about the feature and more about the list of safeguards.
Most of the traps fail silently
Going through the safeguards listed in the commit message one by one shows what failure each one prevents.
- Skip my own comments. The bot's replies show up in the comment list again. Without a filter, it could reply to its own replies.
- Fail closed when the username is missing. If the bot doesn't know my account name, it can't tell which comments are mine. In that case it stops instead of going ahead anyway.
- Write to the ledger only on success. If failures go into the processed-comments ledger too, those comments never get tried again. Recording only successes means failures get retried on the next run.
- Per-account cap. Each account has a limit on how much it can process in a single run.
- Null guard and output guard. These block empty API values and check the model's reply text before it gets posted.
- Multilingual (ko/en/ja). Classification isn't tuned only for Korean comments.
- URL token matched once. The commit note says only this, in one line.
- Live lock. A lock stops two live posting runs from overlapping.
None of these add a feature. Every one of them says: under this condition, don't.
A year name that looks like a slur
One trap only exists because the comments are in Korean. Byeongsin-nyeon (丙申年) is the name of a year in the sexagenary cycle. Its first two syllables are spelled exactly like a common Korean slur. A filter that just matches strings against a profanity list could flag a normal comment about that year as abuse.
In this bot, flagging a comment as abuse means hiding it. A false positive hides a real user's comment. That's a much worse mistake than skipping one reply, so this case is handled separately so it doesn't get flagged.
The real gate was outside the code
One problem was left after all the safeguards went in. Hiding comments and posting replies require the instagram_business_manage_comments scope. No code change gets you that permission.
What would you do here: go live right away, or ship it first in a state where it can't do anything?
I went with the second option.
- DRY by default. Nothing gets posted unless live mode is turned on.
- No-op without the scope. If the permission is missing, the bot doesn't try and fail. It does nothing and exits.
- Scheduled with launchd twice a day, at 13:20 and 20:20.
- Two rounds of safety review with Sonnet. A different model from the one that wrote the code reviewed it for ways it could misfire.
So what shipped is code that's ready to act safely once the permission exists. It doesn't do anything yet.
Self-check
- When your automation's own output (replies, posts) comes back as input, does it recognize that output as its own and skip it?
- Do you write the processing ledger only after a success, or do failures also get marked as done and never come back?
- When a permission or config is missing, is the default to do nothing, or to try anyway?
The honest part
This post records the state at commit time. All I have is the commit message and the file list. I don't know whether the scope was ever granted, whether the bot went live, how accurate the classification was, or how many false positives or misses it had. I don't know whether any false positive came up besides the year-name case, and the commit doesn't record what the two Sonnet review rounds actually caught. I also can't tell from the commit text alone what situation 'URL token matched once' guards against, so I've left it as written.