My benchmark couldn't afford a single share
The gate said 'must beat QQQ day-trading'. That comparator spent all 68 sessions unfilled — a 647-dollar budget shopping for a 661-dollar share. For three months I was benchmarking against idle cash.
Paper-only experiment logs. Charts stay as sourced placeholders until real reports replace them.
The gate said 'must beat QQQ day-trading'. That comparator spent all 68 sessions unfilled — a 647-dollar budget shopping for a 661-dollar share. For three months I was benchmarking against idle cash.
A momentum comparator crushed my model, so I checked whether it deserved its own track. Removing a single symbol from the universe flipped all twelve lookback-grid cells negative.
Pinned commit, quantile channel, look-ahead: I verified the integration down to library source. All correct. And all four symbols still couldn't beat 'tomorrow's close equals today's close'.
I preregistered three literature-backed selection rules over a 50-stock universe and ran the screen once. Two momentum, one reversal. All three lost to just holding the same basket.
The required-sample field sat at null while the bot ran for a month. Nobody had asked. Computing it from the backtest gave 59 years to detect the effect size I had pre-registered.
I shut down an insider-buying paper bot at 32 closed trades. Not because it lost: a sign test at p=0.110 gives me no basis to say it lost. I shut it down because the upper bound of the excess-return interval landed below my round-trip friction.
A paper bot submitted an order every day and bought nothing for eight days. The limit price was frozen at 73.23 for five straight sessions, and the staleness gate I had added five days earlier never fired once — because the feed didn't return zero, it returned an old price that looked fine.
Two curves on the dashboard climbed together. The ledger's actual PnL was negative. The culprit wasn't a bug — it was one line of input validation.
I found an 'edge' three times in one day and killed it three times. The last one died to a simulation of a perfectly efficient market.
I wrote 'no power calculation before launch' into the postmortem of a failed experiment. Then I never audited the bots still running. All four had the same defect, and one of them needed 2.6 years to reach a verdict.
I retired 3 bots. The hard part wasn't 'kill or keep' — the pre-registered gate had already decided that. The hard part was scrubbing the traces left in launchd, dashboards, and disk.
I audited 7 paper trading bots with one sentence: 'Is this profit skill, or just a bull market?' Five died on that question. The absolute returns were almost all lies.
I built nine US-equity timing strategies. All of them lost to SPY on raw CAGR. The only thing that beat it was a one-time static 1.2x leverage — and that isn't alpha.
Exchange-rate param names, a missing sandbox, and candle history that does exist — the things I hit wiring a paper-trading bot to the Toss Securities Open API.
The single practice that changed my paper-trading experiments: writing the kill/keep rule down before the data comes in. Gates decided in advance can't be moved to fit a story.
Two regime-switching strategies passed the idea test and failed the pre-registered out-of-sample Sharpe gate. Why I killed them instead of tuning until they 'worked'.

A paper-trading bot where an LLM makes the call, wrapped in a pre-registered main gate and a separate early-close adoption gate. How I decide whether closing early beats holding.
Practical notes from wiring a paper bot to Polymarket — the Gamma search constraint, the CLOB price-history interval, and why prediction markets are an under-covered niche to build in.

A tour of the paper bots I run — swing bots, a TQQQ infinite-buy experiment, and a trend bot whose most useful finding was that a static leverage beat the timing logic.