I tried the existing tools first
With a pile of apps shipped, each one needed a short video. Hand-editing was not on the table. A generation pipeline was the right shape.
I reviewed MoneyPrinterTurbo and OpenMontage. I ended up building my own: edge-tts for voice, ffmpeg for composition. Narration text and subtitle fragments (subs/s1 through s5, plus a footer) lived as separate files, and a batch script walked the app list.
The first renderer (render_short) produced technically fine videos. Voice played, subtitles appeared, the length was right, and batch mode turned out several at once. But the thing in the center of the frame was the app icon. An icon floats there, narration describes the app, subtitles scroll.
So, a question. Does your promo video show your product, or does it show your logo while a voice describes the product?
Running and worth showing are different things
As a pipeline, version one was a success. From input (app metadata) to output (mp4), no human touched it. Batch produced multiple videos, and no stage failed.
The problem was that the output gave a viewer nothing. An icon is a mark that identifies an app, not a screen that shows it. An icon-only video can say "this app exists" and nothing more. It cannot say "here is what you see when you use it." In short-form, the opening seconds are the whole thing, and I was spending them on a logo.
Fixing it meant rewriting the renderer. Swapping the source from icons to real app screens was not a one-line asset change; it meant redeciding where screens come from and how they are composed. There was no cheap fix.
What would you do?
Keep the working pipeline and increase the volume, or rewrite the renderer because the output is unconvincing?
I redesigned it as a hook renderer
The second renderer (render_hook) starts from a different premise.
- Source: real app screens instead of icons. App screenshots sit inside a phone mockup, so the frame looks like someone using the app.
- Audio: no voice narration. Music and sound effects only. Instead of making people listen to an explanation, make them watch the screen.
- Copy: a separate path for hook copy (hook_copy) and another for gathering source material (hook_harvest).
I did not delete the old path. render_short and render_hook, batch and hook_batch, all still sit side by side. The composition, subtitle, and upload skeleton from version one carried over intact. What changed was what goes on screen.
The app list needed a single source of truth
Once the renderer loops over apps, where the app list comes from becomes the problem. Hardcode the list in each script and it drifts as the count grows.
So I picked one manifest. ootssu.com/apps is the canonical source, with a harvest step and an enrich step feeding a manifest build. The video pipeline and the uploader read the same list.
The uploader (upload_youtube) does not just push video. It fills in the description, the hashtags, and the first comment. The first comment is exactly the kind of step that gets skipped when done by hand, so binding it to the upload stage was safer.
What auto-selection could not do
Screenshots were the remaining problem. Each app needed screens chosen for the phone mockup, and I tried to choose them automatically.
Auto-selection hit a clear limit. The machine could not judge "does this screen explain this app best." On top of that, the screenshot pile had contaminated shots mixed in. I excluded those and picked the rest by hand. Nine final.
Which means one human step remains in an automated pipeline. I did not force automation onto it. A single wrong screenshot throws away an entire video, while picking costs a few minutes per app. In that stretch, the gain from automating is smaller than the risk.
Self-check list
- In the first three seconds of your product video, is it a real product screen or a logo?
- Does the product list your pipeline reads have one canonical source, or a separate copy per script?
- When did you last look with your own eyes at what the automated selection step actually picked?
The honest part
This is a record from the moment the pipeline started running. Whether the hook renderer actually performs better than the icon version, I cannot say with data. I have not checked views, watch time, or install conversion. All I verified is one judgment — that an icon-only video fails to show the app — and that is my judgment, not a measurement.
The nine screenshots were also picked by hand, so a different nine might have led somewhere else. Whether abandoning auto-selection was right will only be clear once video results accumulate. The full journey, architecture, and open items are written down in docs/JOURNAL.md.