A loop where finishing the video wins
My YouTube automation has a loop that's supposed to improve itself. It reads performance data for posted videos, gives more weight to whatever did well, and uses those weights to choose the app and hook for the next video. Until now, the performance metric was average view percentage.
The problem is that view percentage only asks one question: did they watch to the end? A video that plays to the end scores well. Whether anyone shared it or hit subscribe never entered the score.
Is the metric your loop rewards the outcome you actually want, or something nearby that's just easier to measure?
This came out while another coding agent was reviewing what legitimate levers I had left. A reward built only on retention was leaking in two directions. It promoted videos that didn't convert, and it kept promoting whichever app had already won, so everything would drift toward the same video.
Four leaks
First, the reward only looked at viewing. It couldn't tell 'watched to the end' apart from 'did something afterward.'
Second, retention is a single number. I couldn't tell whether a video was abandoned early or sagged in the middle, and the committee prompt that reviews videos was judging without knowing where viewers dropped off.
Third, the weighted rotation had no brake. An app that won got picked more often. More picks meant more samples, which made it more likely to win again, so the rotation could converge on one app.
Fourth, the description links carried no source information. If someone clicked through to a landing page, analytics had no way to split that traffic by channel, app, language, or hook.
What would you do?
With signal this thin, would you fix the reward formula first, or wait until there's enough traffic to measure?
What I changed
I did both. I fixed what I could fix now, and I planted what can't be planted later.
The reward is now:
reward = views × retention × (1 + engagement)
engagement = (shares + 3 × subscribersGained) / views # cappedCounting a subscriber as three times a share is my judgment that subscribing is closer to real conversion. I didn't fit that number to data. The cap is there so a single low-view video can't take over the reward with a few shares.
The retention curve gets read on its own. For each point in the video's progress (elapsedVideoTimeRatio), I pull the share of the audience still watching (audienceWatchRatio). From that, each app gets a diagnosis of early drop-off or mid-video sag, written to hook/retention_hints.json. That file is injected into the committee prompt, so the review step knows where videos are losing people.
The rotation now has an app-level fatigue factor. The more often an app was picked in the recent window, the more its weight shrinks, by FATIGUE^count. This is meant to stop one winning app from taking over the whole fleet.
Description links now carry UTM parameters (source, medium, campaign, content). I haven't built the analytics reader yet. I tagged the links first anyway because videos that are already posted can't be tagged after the fact. I can build the reader later. I can't go back and add tags that should have been there all along.
Tests cover the engagement boost, the cap, and the zero case for the reward; early drop-off, mid sag, and short data for the curve diagnosis; and whether fatigue actually spreads out the picks.
The part that isn't a cheap fix
Most of this commit changes formulas, and formulas are cheap to fix. The real bottleneck is the signal that goes into them. Engagement is a ratio over views, so with small view counts it swings hard on one or two shares. The cap clips the swing but doesn't create any signal. The commit message itself says the analytics reader is deferred 'until traffic exists,' and that tells you where things stand: the measurement end of the loop is still empty.
Self-check
- Does your loop's reward include the actions you actually want (subscribes, clicks, purchases), or is it one easy-to-measure proxy?
- Does your weighted selection have a fatigue or decay term so that one early winner can't take over everything?
- Have you planted tracking that can't be added later (UTM-style source tags) before your measurement tooling exists?
The honest part
I don't know whether these changes made the videos any better. When I made the commit there were no results to compare against. The 3x subscriber weight, the engagement cap, and the fatigue factor were all set by judgment and have never been validated. I haven't confirmed that the curve diagnosis leads to better hooks, or that the committee prompt makes good use of the hints. The UTM tags are in place, but I haven't read anything from them yet.