DarijaBench: Do AI Models Actually Understand Moroccan Darija?
DEV Community

DarijaBench: Do AI Models Actually Understand Moroccan Darija?

DarijaBench: Do AI Models Actually Understand Moroccan Darija? As a Moroccan student, I use AI assistants every day. They are brilliant in English and French - but I kept noticing something: ask them something in Darija (Moroccan Arabic dialect, spoken by 35+ million people), and the confident answers start wobbling. So I decided to stop guessing and start measuring. I built DarijaBench, a 60-item benchmark that tests whether today's frontier models truly understand Moroccan Darija - and ran it on four of them. ๐Ÿ”— Benchmark: https://www.kaggle.com/benchmarks/soufianzaari/darijabench What DarijaBench tests Darija is a low-resource dialect: it is barely present in training data compared to English, French, or even Modern Standard Arabic. DarijaBench probes three practical skills, 20 items each: Darija → French translation - can the model convert everyday Moroccan sentences into French? Sentiment analysis - can it tell whether a Darija text is positive, negative, or neutral? Moroccan cultural knowledge - can it answer questions asked in Darija about Morocco (the capital, couscous ingredients, the 2030 World Cup hosts...)? Each item is graded automatically with keyword checks, and the final score is the average of the three categories. The whole thing was built with the Kaggle Benchmarks SDK, entirely on the free quota. The results Model DarijaBench score Gemini 3.7 Flash 0.95 Claude Sonnet 5 0.95 Claude Opus 4.7 0.95 GPT-5.6 Luna 0.90 What I learned - The top models have mostly cracked basic Darija. A three-way tie at 0.95 was not what I expected - I assumed Darija would still be a weak spot. Frontier labs are clearly doing better on dialectal Arabic than their reputation suggests. - Sentiment is the easiest skill. Gemini scored a perfect 20/20 on sentiment. Emotional tone ("ู‡ุงุฏ ุงู„ู…ุทุนู… ุฑุงุฆุน!" vs "ุงู„ุฎุฏู…ุฉ ุฎุงูŠุจุฉ") seems to transfer across dialects and languages with little friction. - Translation is the hardest. At 85%, Darija→French was the lowest category for Gemini. Darija is full of French loanwords used differently than in French, plus idioms with no literal equivalent - exactly where word-level understanding breaks down. - Small gaps still matter. GPT-5.6 Luna trailing at 0.90 might look close, but on only 60 items that is ~3 more failures - likely concentrated in the cultural questions, where training-data coverage of Moroccan specifics (not just "Arabic") decides the outcome. - Honest caveats. Keyword-based grading is brittle - a correct answer phrased unexpectedly can fail. Sixty items is small. And the three-way tie tells me the benchmark is currently too easy to separate the best models. The next version needs harder items: idioms, heavy code-switching (Darija/French/Spanish mixes are everywhere in the north), and regional variants. Try it yourself The benchmark is public - run more models on it, or fork the task and make it harder. If you speak a low-resource dialect, consider building your own version: the Kaggle Benchmarks SDK makes it surprisingly painless, and measuring is always better than vibing. Built for the DEV x Kaggle Benchmarking Challenge. #kagglechallenge Top comments (0)

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.