Decision models are English-first. I measured them in Japanese, then built one.
What happens to System One decision models outside English: a position prior in Laya, an open Japanese model, a translated benchmark, and what transfers back to English. Decision models had a big September. TypeSafe AI shipped Jev and called it the first "System One" model. Liquid AI followed with d1. Open implementations appeared within days: Laya, AnyJev, kev, Lev. The idea is simple and good. Instead of asking an LLM to write JSON and parsing it, you send text plus typed questions and get back probabilities, with zero output tokens: - noul: yes / no, returns P(yes) - choice: pick one of named options, returns a distribution - score: a position on an ordered scale, returns an expected level and a distribution Almost every public number for these models is in English. I work with Japanese text, so I measured them in Japanese. This post is what I found, what I built, and what I'd tell anyone building one for another language. TL;DR - laya-multilingual never picked the first level of an ordinal question in Japanese: 0 of 300. Same in English: 0 of 290. A label-free check for this is now in Laya. - I built sokudan, an open 310M Japanese decision model. On 300 Japanese business messages: choice accuracy 0.880, score RPS 0.075, bool AUROC 0.844. pip install sokudan , runs on CPU, CUDA or Apple Silicon (MLX). - Trained only on Japanese, its choice transfers to English almost intact; its yes / no does not. - I published a Japanese translation of LocalLLaMA's typed-decisions benchmark, with every case back-translation-checked. 1. The first thing I found: a position prior I wrote 300 synthetic Japanese support messages (bench_ja ) with three unseen schemas: route to one of four departments, rate urgency on three levels, flag churn intent. Then I ran laya-multilingual on them. Choice was fine. Score was not: its RPS was 0.232, worse than always predicting the majority class (0.197). Looking at the predictions, it had chosen the first-listed urgency level 0 times out of 300, while that level was the correct answer for a real share of the messages. The test that made this undeniable is simple and needs no labels. Give every option the same text: level 0: a request level 1: a request level 2: a request A model that reads content, not position, should spread its probability evenly. Any preference is pure presentation. Measure slot 0's log-probability against the mean over slots, and the first-slot rate over all orderings of three real options (1/3 if order doesn't matter). I reported it (laya#131), and the maintainer documented it in the README and model card. Then I upstreamed the check itself: - laya#259 (merged): presentation_checks.py , the identical-option control, first-slot rate and a parity check, so any retrained checkpoint can be tested for the prior. - laya#650 (merged, ships in 0.3.22): fixed states in Japanese, Korean, Hindi and Turkish. The first-slot check fails in all four. - laya#753 (open): the same control for choice. Here laya-multilingual favours slot 0 in all five languages. So the avoidance isn't a general dislike of the first option; it's specific to ordinal questions. I ran the same control on kev and Lev. On score questions, both put less probability on slot 0 than on the other slots when every level carries the same text. For Lev, averaging the score over the presented order and its reverse raises score accuracy by 6.9 points on the English bench (8.7 on the Japanese one) and brings the control close to zero (lev#2, open, off by default). Lesson 1: accuracy alone won't show you this. Measure presentation with a label-free control, in every language you ship. 2. Building a Japanese one sokudan (ๅณๆญ, "snap decision") is my answer to "what if the decision model were built for Japanese from the start". - Backbone: sbintuitions/modernbert-ja-310m (MIT). 314.6M parameters. - State and question are encoded jointly in one sequence, with a marker token per option; a small head reads each marker. Score uses a cumulative-link head, since the number of levels is decided per request. - Training data: synthetic, 21 domains, 4,833 documents, 31,243 labelled (document, question) pairs. One RTX 5090. - Apache-2.0. Weights on Hugging Face. pip install sokudan . On the same 300 Japanese messages, with Lev measured on the same machine: | bench_ja (300) | choice acc | score RPS ↓ | score acc | bool acc | bool AUROC | |---|---|---|---|---|---| | sokudan v0.2 (310M) | 0.880 | 0.075 | 0.817 | 0.780 | 0.844 | | Lev (calibrated) | 0.887 | 0.138 | 0.643 | 0.817 | 0.889 | | laya-multilingual | 0.747 | 0.232 | - | 0.543 | 0.523 | | majority class | 0.380 | 0.197 | - | 0.703 | - | The honest reading: Lev, a much larger model (about 9 GB of VRAM in bf16 in my runs) trained for English, is as good or better than my Japanese one on choice and yes / no. Scale buys a lot of transfer. Where the dedicated model wins is the ordinal question, which is exactly where the position prior hurts. sokudan's own cost is small: about 24 ms per bench item on an RTX 5090, 1.3 GB of VRAM, and it runs on a CPU. import sokudan agent = sokudan.load("GeneLab/sokudan-ja-310m") result = agent.predict( "ๅ ๆใฎ่ซๆฑใงๅใ้้กใไบๅๅผใ่ฝใจใใใฆใใพใใ่ณๆฅใ็ขบ่ชใใ ใใใ", {"department": {"type": "choice", "instructions": "ใใฎๅใๅใใใฏใฉใฎ้จ็ฝฒใๆ ๅฝใในใใ", "criteria": {"่ซๆฑ": "ๆฏๆใใป่ฟ้", "ๆ่ก": "ไธๅ ทๅใป้ๅฎณ", "ๅถๆฅญ": "ๆ้ใปๆฐ่ฆๅฅ็ด", "ใใฎไป": "ไธ่จไปฅๅค"}}}, ) print(result["answers"]["department"]["choice"]) # ่ซๆฑ There's also a server that speaks the same /v1/systemone request and response shape as TypeSafe's public API reference (and Liquid's d1 uses the same shape), so a client written for that format can point at localhost . 3. What transfers back to English I trained sokudan on Japanese only, then ran it on the English version of the bench (290 messages): | English | Japanese | | |---|---|---| | choice acc | 0.872 | 0.880 | | score RPS ↓ | 0.114 | 0.075 | | bool acc | 0.690 (majority 0.683) | 0.780 | Choice carries over almost untouched. Score degrades but stays useful. Yes / no collapses to the majority class. My reading: choosing between named options leans on matching the option text against the state, which a mostly-Japanese backbone can still do in English; a calibrated yes / no needs the model to have learned what "yes" looks like in that language. So: sokudan is a Japanese model. Please don't use it for English yes / no questions. Lesson 2: if you're building for a non-English language, report the transfer table both ways. Which question types survive the language shift is not obvious, and it isn't the same for all three. 4. A Japanese typed-decisions benchmark My bench is one family of synthetic support messages. To test outside it, I translated LocalLLaMA's typed-decisions (Apache-2.0) into Japanese: 1,600 cases across agent traces, customer service, invoices and security incidents. Machine translation needs checking, so every case was back-translated and judged against the source by a second model. Overall match: 0.946. Every case carries its match flag, so you can filter. Dataset: GeneLab/typed-decisions-ja. Zero-shot on the 400 test cases, same scorer for both languages: | Japanese | English | | |---|---|---| | Lev | 0.628 | 0.635 | | Qwen3-8B via AnyJev | 0.625 | 0.648 | | label prior | 0.483 | 0.483 | | sokudan | 0.424 | 0.435 | | laya | 0.348 | 0.353 | Two things. Every model loses at most 0.023 going from English to Japanese, which is decent evidence the translation kept the difficulty. And sokudan is below the label prior here: agent traces and invoices are a different world from the business messages it was trained on. A small dedicated model is dedicated. 5. Things that bit me Option descriptions matter enormously. On the department question, giving each option a short description versus bare labels: choice accuracy 0.880 vs 0.537. Short labels like "billing" or "technical" only mean something once described. Catch-all options are under-chosen. When "Other" is the right answer, sokudan picked it for 11 of 38 cases, though P(Other) still ranks those cases well. Threshold P(Other) instead of taking the argmax. I later found part of the cause in my own training data: 904 rows had a "none of these" option, and it was never the correct answer. The model learned exactly that. Three seeds lie. My first release reported a three-seed mean. Score accuracy across those seeds was 0.663, 0.800 and 0.827; the mean described no actual run, and the shipped weights were the 0.663 one. I moved to 16 seeds per condition, exact permutation tests, and decision rules committed before training. Most ideas I tested afterwards failed that bar. The one that passed was the simplest: v0.2 is a uniform average of eight seeds' weights. The same fix works on one model and not another. Order averaging lifts Lev's score accuracy by 6.9 points. On sokudan I tried three versions of it (at inference, as a training penalty, and by shuffling options in the data); each moved some position checks, but none kept accuracy. Measure the fix on your model; don't import it. If you're building one for your language - Write a small bench in your language with schemas the model has never seen, and keep it out of training. - Run the identical-option control and the first-slot rate per question type. Laya's presentation_checks.py already has fixed states for five languages; adding yours is a few lines. - Report the transfer table to and from English. - Use enough seeds to see the spread before you believe a difference. - Check your training data for options that are always wrong. Links - sokudan: GitHub · Hugging Face · PyPI · Colab - Benchmarks: bench_ja /bench_en in the repo (CC BY 4.0, evaluation only) · typed-decisions-ja - Listed in the Jev Decision Index sokudan is an independent open-source project, not affiliated with TypeSafe AI or Liquid AI. "System One" is TypeSafe AI's name for this category. TypeSafe's terms prohibit using its service or outputs to build similar products, so I have never called Jev, and none of the numbers above are Jev
Comments
No comments yet. Start the discussion.