Automation Pipeline3 min read

My regex worked in English and nowhere else

I built a gate that checks live store copy against what the app really does. It caught things in English, and Japanese and Polish copy walked straight through it.

#python#localization#verification#gotchas#tooling
Concept diagram: a unicode word boundary failing to exist after CJK characters, so the match never fires
A diagram summarising the post.

I built a gate that checks whether live store copy across dozens of apps matches what the apps actually do. If the copy claims iPad support, the gate checks that an iPad build exists.

Tested in English, it caught things. The moment I added other locales, Japanese and Polish copy walked through the whole net.

How many languages has your checking rule been verified in?

Word boundaries were the cause

It was \b. Python's word boundary is decided by what Unicode considers a word character, and by default Han, kana and Hangul all count as word characters.

"・iPad最適化UI"       -> a Han character follows "ipad", so no boundary exists
"・Apple Watchアプリ"  -> same
"Układ … dla iPada"   -> Polish inflection makes "ipada" one word

English copy usually puts a space or a period after a word. It was only ever matching by luck; the boundary logic had never considered another language.

The fix constrains the boundary to ASCII alphanumerics.

def T(*terms):
    return "|".join(rf"(?<![a-z0-9]){t}(?![a-z0-9])" for t in terms)
 
T("ipad[a-z]*")
# matches "iPada" (inflected) and "iPadOS", not substrings like "flipada"

Here is where it splits

Fix that and the next problem arrives: copy that says "does not support" rather than "supports."

In Korean, Japanese, Turkish and Hindi, the negation follows the term. "HealthKit 미사용", "HealthKit非連携". Logic that looks only before the match misses all of them.

Would you scan both sides of the match, or search the whole sentence for negations?

Both have traps. Multi-character negations need a whole-fragment search, but Han negations (不, 未, 非) are single characters, so a whole-fragment search hits unrelated words. Single-character negations had to be searched in a narrow window around the match.

The other two traps

  • Translated feature names need per-language tokens. A term like "complication" becomes a completely different word in each language; a regex looking for the English one never fires. This has its own trap — one Traditional Chinese term means both "widget" and "ordinary tool," so mechanical matching produces false positives.
  • Strip contact details first. A single personal email address inside the copy caused every locale containing it to be flagged as claiming cloud sync. Without removing emails and URLs up front, the check's own output is contaminated.

Self-check

  • Does your regex use \b? Have you tested it against a string with Han or Hangul attached?
  • Does your negation detection look only before the term?
  • Do you strip emails and URLs from text before checking it?

The honest part

Once the gate ran across locales, it found copy selling features that were never built, across several apps and locales. Some of those belong to the same lineage as the three false sentences in a store listing.

Here is the uncomfortable part. Had I verified the regex in English and moved on, those defects would have stayed invisible. Having a gate guaranteed nothing — the gate was green because it could not read, not because there was nothing to find.

Take one regex of yours containing \b and feed it a line with a CJK character attached.

Related