I run three apps built on a realtime voice API. Looking at session records, 106 of 156 had zero audio tokens in both directions.
Tokens minted fine and sessions consumed 32 seconds on average, so this was not a connection failure. Which means nobody opened their mouth.
And these apps required the user to speak first. A realtime session waits for the other side until you send response.create. Voice activity detection cuts turns; it does not open the first one.
Only one of the three apps had that call.
What do you put between a hypothesis and a fix?
The story fit too well
The target user is an adult with speech anxiety. "We were asking the anxious person to break the silence" is a persuasive sentence. There were numbers too: users whose first session was silent averaged 1.79 sessions, against 3.04 for users who exchanged audio.
One response.create inverts the order. I added it to the other two apps and submitted for review.
The next day I looked again
This time I put all three in one table.
| Sessions where user spoke | Sessions where app spoke | |
|---|---|---|
| Greets first | 8 | 20 |
| No greeting (A) | 16 | 16 |
| No greeting (B) | 51 | 51 |
The bottom two rows are exactly equal. The app only spoke after the user did — direct evidence that no first turn exists. That comparison was genuinely useful.
The problem is the top row. The greeting app shows 20 > 8, so it really is speaking first. And its user response rate is the lowest of the three. Eight of 45 sessions (18%) had user speech, against 37% and 39% for the others.
Its median session was shorter, too.
So I had a control group
One of the sibling apps forked from the same codebase already had the feature — for two weeks. The refuting data was in the database at the moment I formed the hypothesis.
I shipped without looking at it.
And the metric itself was wrong
Digging further, audio_out_tokens = 0 is not evidence of silence. Usage is written when a response completes, so a session that ends before the first turn arrives records zero. That is why even the greeting app sat at 63.6% zeros.
The number I read as "106 silent sessions" could not measure the effect of a first turn at all.
A fork in the road
The fix is already live. Roll it back, or leave it?
What would you do?
I left it. Not because the effect is proven, but because it is harmless and the remaining evidence still points that way. What I did change was the metric for judging it: session length distribution and return rate, not tokens. Tokens were never capable of measuring this.
Self-check
- After forming a hypothesis, do you look for data you already hold that could refute it?
- Do you have variants forked from the same codebase? That is a free control group.
- Have you confirmed your success metric can physically observe the change you made?
The honest part
The fix itself is not bad. I still believe greeting first suits this audience.
What was wrong is the order. I shipped before confirming the cause, and when I went to confirm it, I found a refutation instead. It is the mirror image of tests passing while the button did nothing — same root, opposite direction: trouble grows wherever a verification step was skipped.
Before your next fix ships, ask once whether a sample capable of refuting the hypothesis is already in your database.