Have you ever opened a new instrument the day after you built it? Not on build day — the next one.
Yesterday I shipped three of them: a daily time series of follower counts, a ledger of who my bot followed, and one new action. I wrote unit tests. They passed. I committed. I even wrote a post about it.
This morning I looked. All three were broken.
One: dead on its first run
ModuleNotFoundError: No module named 'requests'
LastExitStatus = 256When I registered the scheduled job, I pointed it at the system Python. The dependency lives only in the project's virtualenv. Running it by hand from my shell worked, because my PATH picks a different Python.
Here's the embarrassing part. I had fixed this exact mistake in another repo the previous day. Five judgment scripts had python3 in their usage lines; I changed them to the venv path and wrote "it dies on judgment day" in the commit message. Then that evening I registered a new job with the system Python.
Two: recording wrong values
The follow ledger had 40 entries and only 4 names. And those four looked odd — sh, kim, two or three characters.
My approach had been "take the text that vertically overlaps the button and looks like a username." When I actually dumped the screen, two things came out:
- In the mutual-follow list, the name field really is the username.
- In the suggestions list, the name field is a display name. It holds a nickname in Korean, and the username isn't on that screen at all.
So on the suggestions screen nothing matched, and fragments from neighboring rows did. sh was part of somebody's name.
The worst part is the comment I had written in that very function: "return None if it can't be read — writing a guess means unfollowing the wrong person later. An empty value beats a wrong one." I wrote that, then implemented the opposite.
Three: zero events, and no log line either
The new action never ran all day. The reason is simple: the function guards on being on the right screen and bails out otherwise. And the top of the home screen doesn't have that anchor. There's a story tray; post rows only appear after one scroll.
The day before, I had called this function on a real device and watched it succeed. The screen was already scrolled at the time. My verification skipped the actual entry condition.
What would you fix first here? Adding a scroll retry felt urgent, but that isn't what I fixed first — I made the silent bail-out print why. Zero being indistinguishable from a failure was the bigger problem.
And the dashboard was green the whole time
None of the three showed red on my control board, each for a different reason.
- The dead job still had yesterday's artifact from my manual run, so it looked fresh.
- The ledger with wrong values kept growing. Files grow honestly.
- The zero-event action had no row on the board at all.
Digging further, the board never read the scheduler's exit status. I built this thing on the principle that "exit 0 is not evidence" — and never implemented its inverse, "a nonzero exit is evidence." I had stopped trusting a signal and then stopped looking at it entirely.
Bonus: my tests were writing into production data
Twenty of the ledger's sixty rows came from my own test run. The test file imports the module as-is, so the record path pointed at the real ledger. Those rows have an empty account field, which is why I spent a while this morning asking "why did the bot follow without an account?"
Three things to check
- Do you have a step that opens a new instrument the next day? Build-day checks tend to run under conditions the builder chose.
- When you verify, do you arrive through the real entry path, or do you call the function from a convenient state?
- Does your test suite write to a different path than production? Go check.
The honest part
Every defect here is one I created yesterday. It wasn't really a skill problem — it was an ordering problem. Build, test, commit, move on. The missing step is one line long: read it back a day later.
I already had a "verify after write" rule for external API calls, because a 200 doesn't mean the field was stored. I just never applied that rule to my own instruments. I distrusted the 200 and trusted the green light.
Pick one log or metric you added recently and open the file itself. "Rows are growing" and "the values are right" are different claims.