I run three apps you talk to out loud. Across 218 sessions that ended cleanly, 143 never delivered a single word of user audio to the server. That is 66%.
Per user it is sharper: of 130 people, 87 never spoke once.
When I first saw that, one thought arrived. The microphone is not being captured.
How do you tell "the user did not do it" apart from "the system did not receive it"?
Ruling things out in code
Permission denial. Audio start runs before the server session is created, and a denial stops there without ever calling the backend. A test pins this: "a microphone failure creates no server session at all." So every silent session in the database belongs to a user who granted the permission. Ruled out.
The echo gate. There is a second line of defence against speaker bleed reaching the mic. If it swallowed user speech, this is exactly what it would look like. Reading it: it only engages within 0.4 seconds of the last audible playback frame and passes everything through otherwise. In the two apps that do not greet first, there is no playback when the user speaks first, so the gate is transparent. Ruled out.
Connection failure. Token mints average 1.0. The socket opened. Ruled out.
At that point I thought the only way forward was new instrumentation. Logging the capture buffer count would separate "the mic delivered but it never left" from "the user said nothing." That value was in fact already collected — but deliberately written to logs instead of the database, to avoid changing what an anonymous-auth app stores and therefore its privacy labels. And the logs had aged out.
Then I looked at the clock
While drafting the instrumentation plan, I split session length by category on a hunch.
| Category | Count | Median seconds consumed |
|---|---|---|
| Nobody spoke | 129 (59%) | 4–12s |
| App spoke only | 14 | 12–180s |
| Real conversation | 73 | 30–34s |
Four seconds.
First-response latency alone is 1–3 seconds. No conversation can physically happen inside four seconds. If audio were genuinely broken, a user would try, wait for a reaction, try again — fifteen to thirty seconds, minimum.
The instrumentation was unnecessary. A column I already had held the answer.
The control group settled it
Only one of the three apps speaks first. The other two require the user to break the silence.
Even in the app that greets you, "nobody spoke" came to 25 sessions with a median of 4 seconds — closed before the greeting could arrive. Meanwhile the 12 sessions where the greeting did land ran to a 24-second median, and the user still never answered.
So the problem was not "it does not speak first." That was a hypothesis I had formed days earlier and already shipped a fix for, and the only control group I had disproved it.
Self-check
- Have you looked at dwell time for your failure samples? Is it shorter than the minimum time your system needs to respond?
- Is the column that separates "the feature broke" from "there was no opportunity" already in your table?
- Before adding instrumentation, do you spend five minutes checking whether existing columns in combination can answer it?
The honest part
This is not good news. A technical defect can be fixed; people opening a call screen and closing it in four seconds is not a code problem. It sits in the same place as more installs and still zero revenue — growing acquisition does not get you past this gap.
The diagnosis did get honest, though. Two more days chasing an audio bug that was not there would have cost exactly that.
Run one query for the median dwell time of your failure samples. If it is shorter than your system's response time, the bug you are looking for may not exist.