Is regex enough? I tested span-01 on mixed-language text
DEV Community

Is regex enough? I tested span-01 on mixed-language text

Have you ever exported a script for text-to-speech and found, later, that a single English word had slipped into it? I have - and the synthesized voice mangled that spot badly enough that I had to redo the take. All I actually want is to check "is there any English mixed into this Japanese script?" That turns out to be surprisingly annoying. Checking everything by hand is not realistic, so the standard move is to filter with a regular expression. Then I noticed that a service called respan ships models "built for decisions" - the official site is respan.ai. So I gave it a try, thinking it might take some of the pain out of that problem.

Code and measured data: github.com/sunnydachs/span01-eval

The verdict first

On boundary cases, the model beats the regex by a wide margin. In sentences that mix in brand names, URLs, personal names or acronyms, the regex flags every one of them as a defect. Conversely, on my real scripts, the regex was enough. That yardstick is not fair, though, and I explain why below. So the practical shape turned out to be "regex as the first stage, the model as the second" - it is not a case of picking one over the other.

What this thing actually is

It is an unusual kind of model that respan provides. The ones I looked at were span-01 and span-01-lite. It is not a chat model that generates text. It is a model that returns a probability for whether something matches a condition. You describe what you want to detect in plain language, send it, and get back a probability between 0 and 1. Think of it as building a classifier with a prompt instead of with training data. The rule for the decision lives in prose, not in code. As I understand it, that is the selling point of this model.

The problem I hoped it would solve

The problem from the top of this post fits that shape exactly. What I want to detect is "is English mixed into this Japanese script?" The incumbent method was a regex, and it genuinely works well. But there are places where it breaks quietly:

  • A single English word attached directly to Japanese text - missed.
  • The opposite: brand names, personal names and code fragments - falsely flagged.

I figured a model that lets me define both cases in plain language could cut down the misses and the false positives at once.

What I actually tried

I measured it in a comparable way rather than by impression. Two kinds of data went in:

  1. A synthetic suite of 23 boundary sentences I wrote myself: one case each for English words, English sentences, katakana words, brand names, personal names, URLs, acronyms, code fragments and role nouns.
  2. Real Japanese scripts, about 200 of them - every one that contains Latin characters, plus a sample of the ones that do not.

These are the three detectors I compared:

[ Japanese script ]
|
+-- regex ---------- deterministic, instant, free
+-- decision model - rule in prose -> probability
+-- general chat --- asked to answer YES/NO

176 API calls in total.

Result: on boundary cases, the model wins

First, the synthetic suite of 23. F1 rolls misses and false positives into one number; 1.0 is perfect.

Detector F1
Regex (patterns built for this job) 0.52
Regex (catch every 2+ Latin letters) 0.62
Decision model (vague instruction) 0.67
Decision model (exclusions stated) 0.82
Decision model (instruction written out fully) 0.93
General-purpose chat model 1.00

The regex loses by a lot. The reason is plain: the regex produced 9 false positives. iPhone, YouTube, Instagram, URL, personal names - and OK, AI and DX trip it too. All of those are ordinary things to see in a Japanese script, and none of them is a defect. There were also 2 misses, from English sentences not attached to Japanese text. "Deciding where the boundary is" was not a job the regex was suited to.

Result: on real scripts, the regex was enough

So what happened on the ~200 real items? Here the result flipped.

Detector Detected
Regex 22 of 22
Decision model 21 of 22

But this yardstick is not fair. "Ground truth" here is literally "does it contain Latin characters". I collected the items with no Latin characters and defined them as the negative class, so any program that searches for Latin characters scores perfectly by construction. The regex is not strong; I was grading it with the regex's own definition. The real contamination was exclusively "an English word attached directly to Japanese", and for that shape four lines of regex catch everything - a fact, but not evidence that the regex is the better tool.

The most surprising part

What was interesting was not the accuracy but how much the instruction wording matters. For the same sentence - Japanese with one English word mixed in - the result flips on nothing but how the instruction is written.

instruction: "flag English words"
ใ“ใฎmethodใฏใ‚„ใฐใ„ -> probability 0.06 missed

instruction: "count even a single English word"
ใ“ใฎmethodใฏใ‚„ใฐใ„ -> probability 0.81 flagged

(The test string is the literal case from the repo: Japanese with the English word "method" glued into it.)

There were also cases where the model ignored the instruction. Even when I wrote "detect all Latin characters", brand names and URLs kept being excluded. The model carries a built-in prior along the lines of "this kind of thing is probably fine to leave alone", and an instruction cannot fully override it. Being able to define things in plain language cuts both ways. The more you write, the better it gets. Write too little and it misses quietly.

One correction about thresholds

Because the model returns a probability, where to draw the line is a real question. My first instinct was to write "use two thresholds, 0.15 and 0.85" - measuring it showed that was wrong.

threshold F1 false positives
0.15 0.86 7 <- worst
0.50 0.89 4
0.70 0.92 2
0.85 0.83 2 <- more misses

Drop it to 0.15 and the model starts picking up OK and code fragments again - the very things it had supposedly learned to leave alone. 0.5 was good enough. There is no reason to lower it.

Honest limitations

  • The synthetic suite is only 23 items, each measured once. The gap between 0.93 and 1.00 is one case. It is not grounds for a ranking.
  • The instruction was chosen while looking at those same 23 items, so there is no held-out set.
  • The ground-truth labels are my own judgement, and I corrected one after the runs.
  • Before publishing, I had five independent AI models recompute the numbers and fixed the errors they found.

Summary: how to use it

Put the regex in the first stage, and use the model as a second stage to reduce boundary false positives. As a flow:

flowchart TD
    A[Script] --> B{Regex hits?}
    B -->|yes| E[Fix it]
    B -->|no| C[Send to decision model]
    C --> D{Probability 0.5 or more?}
    D -->|yes| E
    D -->|no| F[Pass through]

This shape has three advantages:

  • Determinism is guaranteed at the regex stage.
  • The model is asked to do exactly one thing: not to false-positive on brand names.
  • Swapping the model out does not break the overall decision.

Conversely, concluding "the regex is enough" from real data alone would have been dangerous. My scripts happened to contain only "adjacent" contamination, so the moment my writing produces another shape, the regex stage falls over silently. And a purpose-built model was not actually necessary: a general-purpose chat model can do the same job, with retries. Still, "you can write the exclusion rules in plain language" seems likely to pay off when you need conditions that are awkward to express as a regex.

Code and measured data: github.com/sunnydachs/span01-eval

This is a personal OSS project, so there is no warranty. Use it at your own risk. Bug reports and improvement ideas are welcome as issues.

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.