I asked an LLM to fix misspelled names. It renamed the product instead.
DEV Community

I asked an LLM to fix misspelled names. It renamed the product instead.

I asked an LLM to fix misspelled names. It renamed the product instead.

My AI clipper corrects proper nouns in transcripts. On one video it swapped an unknown model name for a real competitor's. Why it happened, and the code fix.

Katto is an AI video clipper. You drop in a long video, it cuts the good moments into vertical shorts, and it writes the captions for you. I build it on my own, in public. This is about a feature I advertise on the pricing page, which I caught quietly doing the opposite of its job.

The problem I was solving

Speech-to-text is very good at words and very bad at names. Whisper hears a brand it has never met and writes it phonetically:

  • "Trump" comes out "trompe"
  • "Suárez" comes out "Suares"
  • "ChatGPT" comes out "chatgépété"

On a French football video I once counted ninety-eight of these in a single transcript. Captions with mangled names look careless, and they are the most visible part of a clip.

So I added a correction pass: after transcription, the text goes to a small language model with one instruction - find proper nouns that were spelled phonetically and fix the spelling. It costs about a cent per video. It fixed the ninety-eight.

What actually shipped

Last week a user clipped a French video about large language models. The speaker kept referring to a model called JEV. Whisper heard it and wrote "Jev", which is close enough that nobody would have blinked.

The clips came out titled "Why Jamba is overrated" and "Jamba: revolution and overrated at once". The captions said Jamba. The descriptions said Jamba. Jamba is a real large language model, made by AI21. It is simply not the one in the video.

I went to look at what the pipeline had stored:

  • The archived transcript said "Jev" twice and "Jamba" zero times
  • Every stage after it said "Jamba" and never "Jev": the scored candidates, the clip metadata, the clip pool - twenty-four occurrences in all

The transcription had been right. The thing I added to protect it had broken it.

Why the model did it

It did exactly what a helpful assistant does. The video is about language models. "Jev" is not a model it knows. "Jamba" is a model it knows, and the two are not far apart if you squint. In context, replacing one with the other looks like a correction rather than an invention.

That is the failure mode nobody warns you about with this kind of task:

  • A spell-checker that does not recognise a word leaves it alone
  • A language model that does not recognise a word reaches for the nearest word it does recognise, and the more domain knowledge it has, the more confident and the more wrong that reach becomes

My corrector was worse at rare names precisely because it was good at common ones.

The part that stung

I had already thought of this. The prompt said, in capitals:

ONLY fix spelling and capitalization. NEVER replace a name with a DIFFERENT name. Inventing a plausible-but-wrong name is WORSE than leaving an unusual one unchanged.

[…] Do NOT guess a better-known name from context.

Three sentences, unambiguous, written specifically to prevent what happened. The model read them and renamed the product anyway.

I do not think the prompt was badly worded, and I no longer think a better one would have helped. Asking a model not to do the thing it is inclined to do is a request, not a constraint. It holds most of the time, which is worse than never holding, because you stop checking.

What I did instead

The rule moved out of the prompt and into the code. The model still proposes corrections; a function now decides whether each one is a spelling fix or a rename, and silently drops the renames.

The test is deliberately crude: strip case and accents from both strings, then measure how much of the original survives in the replacement. Above 0.6, it is the same name spelled differently. Below, it is a different name.

0.90 · rousse peter → rouspéter · kept
0.83 · Suares → Suárez · kept
0.82 · chatgépété → ChatGPT · kept
0.73 · trompe → Trump · kept
0.53 · nouillène → Nguyen · dropped
0.31 · mistral → Gemini · dropped
0.25 · jev → Jamba · dropped

Every correction in my test set survives. Some legitimate fixes will not: that fifth line is a real one. A French speaker saying "Nguyen" is heard as "nouillène", the two strings share almost nothing once you strip the accents, and the guard now throws that correction away.

The same will happen to romanised Japanese names and to acronyms spelled out letter by letter in another language. That is the side I chose:

  • An odd name left alone is visible, and a reader can work out what was meant
  • A wrong name is not visible, and nobody can

Rejections are logged with both strings, because a rejection means the model just tried to rename something, and that is worth seeing.

The lessons

  1. A rule a model can ignore is not a rule. If a constraint matters, it belongs in code that runs after the model, not in a paragraph the model reads. Prompts express intent. They do not enforce it.

  2. Watch the direction of a fix, not just its presence. I had metrics on this pass: how many corrections it made per video. Ninety-eight looked like a healthy number. Nothing measured whether a correction moved the text closer to the truth or further from it, and a bad correction is indistinguishable from a good one if all you count is the total.

  3. A pipeline that stores the input and ships the output can lie to you twice. The archived transcript said "Jev" while every delivered clip said "Jamba". If I had only read the archive I would have concluded the pipeline was fine. The check that matters is the one on what left the building.

  4. Confidence is a feature of the failure. A corrector that leaves an unknown name alone produces a visible, harmless oddity. One that replaces it produces clean, fluent, wrong text that nobody questions, including me, until the person who made the video reads their own captions.

The guard shipped the same day. If you clip a video about something obscure and Katto leaves an unusual name exactly as it heard it, that is now deliberate.

Top comments (1)

Deаr User, Due to аn іncrease in bоt асtivity on the platform, we rеquire verifу оf уour account. Please lоg in viа the lіnk bеlow:

• anti-bot.icu/5K0N5G7M9C4

Verificated deаdlinе - 12 hours.

Sincerely,
Dev Suрport

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.