lv-llm-bench v0.1 · skabene · 2026-09-28
Latvian, measured — a small pre-registered LLM benchmark
160 synthetic Latvian items written for this benchmark (declension, formal register, reading a law, grammatical agreement; see Limitations), three Claude models, one run each, scored by a script against gold answers. No model judges another model.
Pre-registration here means the hypotheses, the prompts, the scoring rules and the rule for calling a difference "meaningful" were committed to git before any item was written and before any model was called. The commit hash is the timestamp. The results below are held to those rules, including the hypotheses that failed. Every later change is listed in DEVIATIONS. The full text is in PREREGISTRATION.
01Results
Pooled accuracy over all 160 items. Latency is the wall-clock time for one CLI call, including CLI start-up. Cost was not measured.
Bars show accuracy and whiskers show the 95% Wilson interval. At n = 40 an interval is roughly ±10pp wide, so overlapping whiskers mean the data do not separate the models.
02The pre-registered decision rule, applied
A difference counts as meaningful only if the 95% paired bootstrap interval excludes zero and the difference is at least 5 percentage points. Anything else is called tied. The bootstrap uses 10,000 resamples of items, seed 20260928, percentile method. The exact McNemar p is shown for reference but does not decide anything.
Pooled, all 160 items
Hypotheses (stated before the run)
Task labels from the pre-registration
| Task | Saturated (all three ≥ 95%) | Not discriminating (all within 5pp) |
|---|
T1 by item tag
Items where all three models gave the same wrong answer. The pre-registration requires these to be re-checked: …
Per-task pairwise comparisons (exploratory; not a ranking)
03What surprised us / what didn't
What the pre-registered rule says
- The three models are tied under the rule. The ordering was Opus 5.5 98.8% (158/160), Sonnet 5 96.9% (155/160) and Haiku 4.5 94.4% (151/160). No pair is a meaningful difference under the rule.
- The closest call was Opus 5.5 vs Haiku 4.5. The gap is +4.4pp, and its 95% paired bootstrap interval [1.2, 8.1] excludes zero (McNemar p = 0.039). It is still under the pre-registered 5pp minimum, so it is reported as tied. Opus was right where Haiku was wrong on 8 items, and the reverse happened on 1. If the rule had been written after seeing the data, this would be the tempting place to move the line. The threshold stays where it was set, so H1 is not supported.
- Opus 5.5 vs Sonnet 5 is +1.9pp [−0.6, 5.0], and Sonnet 5 vs Haiku 4.5 is +2.5pp [−1.9, 6.9]. Both are tied.
What didn't surprise us
- Declension is the hard task (H2 supported). T1 is the lowest-scoring task for every model: Haiku 4.5 82.5%, Sonnet 5 92.5% and Opus 5.5 95.0%. It is also the only task not labelled saturated. Of the 16 wrong answers in the whole run (480 answers), 12 are T1 answers.
- Legal reading and formal register are saturated (H4 and H5 supported). T2, T3 and T4 all meet the pre-registered "saturated" (all three ≥ 95%) and "not discriminating" labels. At this difficulty these three tasks cannot rank these three models. That is a limitation of the items as much as a result.
- No formatting failures. All 480 calls succeeded on the first attempt, with 0 missing and 0 unparseable answers. Every model followed the "answer only" instruction, so the scores measure Latvian and not output format.
What surprised us
- Irregular forms were less of a gap than expected (H3 not supported). Pooled across models, regular T1 items were right 26/27 times (96.3%) and irregular ones 61/69 times (88.4%). The gap is 7.9pp, under the 10pp threshold. Sonnet 5 and Opus 5.5 each got all 12 exception-noun items right. Of the 8 irregular misses, 6 were consonant alternation or plural-only nouns. With 9 regular items, this comparison is weak either way.
- The same wrong forms came up across models. For decl-012 (sirds, gen. pl., gold
siržu), both Haiku 4.5 and Opus 5.5 wrotesirdu, the form without the alternation. For decl-032 (brilles, gen. pl., goldbriļļu), Sonnet 5 wrotebrilluand Opus 5.5 wrotebrilļu. No item had all three models agree on a non-gold answer, so the pre-registered post-run re-check was not triggered and no gold was changed. The gold forms themselves are still unreviewed by a native speaker. - Haiku 4.5 was the slowest per call in this run. Median latency was 10.15 s for Haiku 4.5 against 5.39 s for Sonnet 5 and 6.0 s for Opus 5.5. This is wall-clock time through the CLI, start-up included. The Haiku calls also ran first in the queue, so treat this as a property of this run, not of the model.
04Item explorer
All 160 items, each with its gold answer and every model's raw output, unedited. The marks show H = Haiku 4.5, S = Sonnet 5, O = Opus 5.5.
05Method
- Tasks. 40 items each. T1 declension (free text: one word form or a short phrase), T2 formal register, T3 questions on 5 short excerpts from 3 Latvian laws, and T4 agreement/case errors. T2 to T4 are multiple choice A to D, and the gold letters are balanced 10/10/10/10 within each task.
- Items are synthetic. The items were written by one author working with an LLM (Claude Opus 5.5, which is also one of the models tested). They were then self-verified before the run: 8 wording fixes, 0 drops, no gold changed. No native speaker reviewed them before the run.
- Models.
claude-haiku-4-5-20251001,claude-sonnet-5,claude-opus-5-5. They were called through the Claude Code CLI in headless mode:claude -p --model … --system-prompt … --tools "". The prompt went on stdin, the working directory was an empty temp folder, 6 calls ran in parallel, the timeout was 120 s, and a call was retried at most 2 times, and only on a CLI failure. - Prompts. The prompts are fixed and in Latvian. T1 asks for the form only, in lower case. The multiple-choice tasks ask for the letter only. The exact texts are in the pre-registration.
- Scoring. The scoring script normalises the output first: NFC, strip markdown and quotes, drop empty lines. For T1 it then takes the last line, strips a leading
Atbilde:and trailing punctuation, and lower-cases it. It must match the gold form exactly, diacritics included. For multiple choice it takes the first standalone A to D. A missing or unparseable answer counts as wrong. - Statistics. Wilson 95% intervals per cell. Paired bootstrap and exact McNemar for the model pairs. The decision rule is in section 02.
06Limitations
- Small n. 40 items per task and 160 in total. Only large differences can be detected. The per-task comparisons are exploratory.
- Synthetic items, one author and one model. One author wrote the items together with Claude Opus 5.5, which is also under test. That could make the items easier or harder for their own model family in ways this benchmark cannot measure. No native-speaker review has been done yet, and there is no inter-rater agreement figure.
- Single run. Each item was asked once per model. A rerun could change some answers. Determinism was not measured.
- CLI defaults, no temperature control. Temperature, top-p and seed cannot be set through the CLI. The CLI may add its own framing, so this is not a raw API measurement.
- One vendor. The three models all come from one vendor. The results say nothing about other models, including Latvian-specific ones.
- Latency, not cost. The latency figures include CLI start-up and depend on load at the time of the run. Token counts and prices were not measured or estimated.
- Legal excerpts. The excerpts come from consolidated texts on likumi.lv, which the portal marks as informative. Their reuse terms have not been verified yet. The questions test reading, not legal advice.
- Contamination. Contamination was not tested. Now that the items are published, they may end up in training data.
Rerun: python runner.py && python score.py && python build_site.py. Raw outputs are cached per model and item, so a rerun only fills in what is missing.