Category

AI-Assisted Dev

Practical workflows for delegating, reviewing, and constraining AI coding agents.

All posts

25
On the left, a verification log reading 'five reviewed, five approved'; on the right, the same log reread as 'zero rejections'
AI-Assisted Dev5 min read

My Review Committee Had Never Once Said No

I added an LLM review stage to an unattended publishing pipeline and committed it as verified. The verification: five Japanese posts, all approved. Here is what an all-approve result proves about a gate, and what it does not.

#llm-judge#content-pipeline#automation#verification
Concept diagram: intended text and typed text differing while the pipeline carries an ok badge
AI-Assisted Dev5 min read

The automation silently typed different Korean syllables

Typing Korean through the input action lands some syllables as different characters. The tool reports success, and the corrupted characters appear in the returned log exactly as they landed. Reading the log cannot catch it.

#agents#automation#verification#gotchas
Three clocks in a row labelled version created, build uploaded and submitted for review, each showing a different time. The leftmost is enlarged with a speech bubble reading 'four days ago!?', while an arrow from the rightmost clock marks the real wait of 20 hours.
AI-Assisted Dev4 min read

I Reported "Waiting Four Days." It Had Been 20 Hours

I ran a script that prints how long each app has been waiting in review and one app looked far older than the rest. That date wasn't the submission date, it was the version creation date. The real wait was 20 hours — and I had already attached the word 'unusual' to it.

#debugging#gotchas#automation#agents
Left panel shows an audit reporting contrast, overflow and page errors all passing. Right panel shows the same screen where cream cards turned khaki and a black scrim became a brown film.
AI-Assisted Dev6 min read

Every Contrast Check Passed — The Screen Was Dead

I recolored 26 apps automatically and ran a WCAG contrast audit across every route. Zero failures. Then a user said the text was barely visible. What the audit never measured was not contrast — it was which color was paper and which was ink.

#accessibility#design-system#first-principles#reality-check
Concept diagram: a prompt forks into an execute path (printing step OK lines) and an explore path (brainstorm -> plan doc -> exit). A guard header blocks the explore fork, but for analysis work that bar goes semi-transparent.
AI-Assisted Dev4 min read

The agent wrote a plan instead of working

I autonomously delegated mechanical tasks to a CLI agent, and instead of working it opened a web mockup or wrote a plan document and exited. The agent hadn't failed — it was doing a different job: an auto-loaded skillset read 'build something' as brainstorming. But the guard that fixes this is useless for analysis work.

#ai-agents#claude-code#prompting#reality-check