Shipping & Infra4 min read

'User cancelled' came back in two tenths of a second

Hundreds of login failures were labelled 'cancelled'. I stopped trusting the label and measured elapsed time — and found two different realities behind one name.

#ios#retention#debugging#verification#reality-check
Concept diagram: the same 'cancelled' reason measuring 0.2s on one platform and 6.5s on another, meaning different things
A diagram summarising the post.

I pulled 28 days of login failure instrumentation in full. Hundreds of entries carried the reason "the user cancelled sign-in."

Instead of trusting that label, I measured the time from tap to result.

What is the evidence behind the reason written in your failure logs?

0.2 seconds is not human time

Successful logins had a median of 7.8 seconds on every platform — the sheet appearing, an account being chosen, approval being granted. Real human time.

On one platform, "cancelled" had a median of 0.2 seconds.

Nobody opens and dismisses a sheet in 0.2 seconds. The sheet never appeared.

Platform Reason Count Median elapsed
A cancelled 318 0.2s
A unknown error 138 0.0s
A success 94 7.8s
B cancelled 40 6.5s
B timeout 17 90.1s
B success 119 7.8s

On the other platform "cancelled" has a median of 6.5 seconds — that is a real cancel. One reason code was pointing at two different phenomena.

A fallback was manufacturing an empty window

The cause was the line that picks the window to anchor the auth sheet to.

.first(where: { $0.isKeyWindow }) ?? ASPresentationAnchor()

That fallback does not mean "no anchor." It constructs a fresh window attached to nothing and hands it over.

The auth system cannot present anything in that window and finishes immediately as "cancelled." The call site records it verbatim as "the user cancelled."

A system failure was wearing a user's decision as a disguise.

Here is where it splits

This is the same dataset as the investigation into half my installs leaving at the login screen. That post was about a structural problem — the server already allowed anonymous reads. This is a different mechanism found in the same data.

Across 28 days, of 97 users who failed login, 51 retried and succeeded and 46 left permanently. Only three of those 46 had ever used the app; the other 43 never got past the first screen. Against new arrivals in that period, 17% vanished at the door.

Looking at "318 cancellations," what would you do — rewrite the login copy, or distrust the label?

Rewriting the copy would have left all 318 in place. Nobody had cancelled.

An interesting side finding

Users who tried two sign-in methods alternately succeeded less often than those who tried one — 2.2 taps and 69% success versus 17.5 taps and 23% success.

They were not stuck for lack of an alternative. Every method they tried failed to open. "Offer more options" is precisely the wrong prescription here.

Self-check

  • Is there an elapsed time next to your failure reasons? Without it there is no way to test the label.
  • Are you assuming a reason code means the same phenomenon on every platform?
  • Is a fallback handing over an empty object instead of nothing?

The honest part

  • One reason code showed 0.0 seconds, and that is a different path from the anchor bug. One sign-in method presents its own UI and never touches this anchor logic. Likely a device-side setting, but I did not confirm the cause.
  • The instrumentation had a hole. The "tapped a sign-in method" event records which method; the "login failed" event did not. I could never split those hundreds of failures by method. Instrumentation is improved now; past data cannot be recovered.
  • The 90-second delay on the other platform is the auth system's own timeout ceiling, and I did not touch it this time.

Take one reason from your failure logs and compute its median elapsed time. If it is a time no human could produce, the label is lying.

Related