Automation Pipeline3 min read

My category ranking had its own definition of good

On my app-download channels, category performance was scored by views times retention. The same repo already had a shared reward function that also multiplied in engagement. The fix was a few lines, and it showed me the system had been running on two definitions of a good video.

#metrics#reward-function#llm-prompt#shorts#consistency
Side-by-side diagram: a category ranking using only views times retention next to a shared reward function that also multiplies in engagement
The ranking that picked what to make next was the only part ignoring shares and subscriptions.

The score was being computed

My short-form channels that aim at app downloads have a step that scores each content category. The resulting ranking goes into an LLM meeting prompt that picks upcoming topics, so it shapes which categories get more videos.

That score was views × retention, with two constants clamping retention to a cap and a floor. It ran fine and produced a ranking every time. Nothing was broken.

The catch was that the same repo already had a shared reward function. The feedback loop uses it to score videos as views × retention × engagement (shares, subscriptions and similar responses). Everything else judged a good video that way. The one step deciding what to make next ignored engagement entirely.

How many places in your system compute "performance," and have you checked whether they all use the same formula?

Why it stayed invisible

Both formulas looked reasonable. Views and retention are obviously performance metrics. There were no errors and no absurd rankings. So nothing showed that this ranking judged categories that lead to shares or subscriptions by the same standard as categories that just get seen.

On an app-download channel that difference matters. A view that leaves a response behind is closer to the outcome I want than one that scrolls past. The meeting that picked the next topics never received that signal.

What would you do? Add an engagement term to the ranking formula, or reuse the function that already exists?

What I did

I reused it. Category performance now scores through the shared reward function, and the two retention cap/floor constants, now unused, are gone. One file, six lines added, seven removed.

Adding a separate engagement term to the ranking would still have left two formulas, and the next edit to only one of them would split them again. With one function, what the feedback loop calls good and what the ranking says to make more of follow the same standard.

Self-check

  • If you search the repo for every place that computes performance, reward or score, is there really only one formula?
  • Does the step that chooses the next action (prompt, priority queue, scheduler) get the same metric as your learning or feedback loop?
  • If a constant (cap, floor, weight) exists in only one formula, can you explain why it lives only there?

The honest part

This commit makes the formulas match. It does not show that results got better. The record says nothing about which categories actually moved after the change, or whether downloads or subscriptions changed, and I don't know yet either. This commit alone also doesn't confirm that the shared function handles the removed retention cap and floor the same way. For categories with little engagement data, one extra multiplication could swing the ranking a lot. What I can say for sure is that the system had two definitions of good, and now it has one.

Related