Tools & Dev Environment4 min read

I read 19% of sessions as silent. The metric could not see itself.

I counted zero audio tokens across 271 sessions in three voice apps and went hunting for an audio bug. Halfway in, I found the number was a tautology, not an observation.

#analytics#verification#debugging#gotchas
Concept diagram: tokens are only written at session-end, so any session without that call is always zero
A diagram summarising the post.

While going through session records for three voice conversation apps, one number stood out. Of 271 sessions, 51 — 19% — had zero audio tokens in both directions. Connected, but nothing said either way.

Token mint counts averaged 1.0, so the socket opened fine. That leaves two options: the microphone never delivered, or the audio stream never reached the server. I started digging into the audio path.

When your dashboard shows a zero, does it actually mean "this did not happen"?

Three hours of digging

I read the echo gate. It only engages during playback and passes everything through otherwise. I read the microphone permission path. A denial throws, and that throw reaches the screen. There was even a capture watchdog: if the buffer count is still zero after three seconds, it restarts audio once.

All of it was fine. So I went deeper and looked at where the server writes token counts.

session-end:   audio_in_tokens: Number(tok.audio_in ?? 0)
heartbeat:     (no code that receives tokens)

That is where I stopped.

What those 51 sessions actually were

They were all in the abandoned state. And when I checked what abandoned means in the code, it turned out to be a session where the client never called session-end — it gets swept up later, when the next session starts.

Tokens are only written by that end call.

So an abandoned session has zero tokens by definition. It does not matter whether audio flowed. What I read as "19% were silent" was not an observation but a tautology: sessions that skipped the end call have nothing in the field written by the end call. There is no information in that sentence.

A fork in the road

There were two ways forward here. One: drop those 51 and recount with what is left. Two: add instrumentation and look again in two weeks.

Which would you pick?

I picked the first. Keeping only the 218 sessions that ended properly, the silent share was 66% — not 19%. The number got bigger, and only then could I ask the real question.

A genuine bug fell out of it

The structure had a real problem. Usage for any session that skips the end call is never recorded. Minutes are settled separately by the server, but tokens stay at zero. Cost was being under-counted by exactly that much.

The fix is simple: carry cumulative usage on the heartbeat that already fires every 15 seconds, and have the server take the larger value per field.

mergeTokenUsage(current, incoming) -> per-field max

It has to be max because a reconnecting app restarts its accumulator at zero. Take the incoming value literally and you erase what was already recorded. Consumed seconds were already protected this way — tokens just sat outside that guard.

It also has to work with older apps that send no tokens at all, so a missing value leaves the stored one untouched. The server ships before the app does.

Self-check

  • Do you have metrics written only on one code path? What does that field read as for samples that never took the path?
  • Are "zero" and "not measured" sharing the same column?
  • Have you actually read the code that assigns your state labels (abandoned, failed, timeout) — who sets them, and when?

The honest part

I did not walk into this because I was unaware of the trap. I had already written down, in this same project, "before judging, check what the metric cannot see." Days later I took a number of exactly that kind at face value.

This is the same family as a test run that executed nothing and still went green and an empty key that passed every gate. The number was not lying. I asked it a question it had no way to answer.

Pick the column with the most zeros on your dashboard and grep, just once, for which function writes it. I did that grep three hours late.

Related